Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Build a mini autograd

Be able to implement a scalar autograd engine (micrograd style) and train a network with it.

Prerequisites

Intuition

An autograd engine needs surprisingly little:

  1. A Value class that carries a number, its gradient, its «children» and a local _backward function.
  2. Every operation (+, ·, tanh …) creates a new Value, connects it to its operands and defines how the gradient is to be propagated backwards through that particular operation.
  3. backward() sorts the graph topologically and runs _backward in reverse order.

That is all. PyTorch does the same thing, but on tensors and with a thousand optimisations.

The key insight: gradients are accumulated (+=). A variable used twice gets a contribution from both paths — that is the chain rule's sum rule.

Code

import math

class Value:
    def __init__(self, data, children=(), op=""):
        self.data, self.grad = float(data), 0.0
        self._backward = lambda: None
        self._prev, self._op = set(children), op

    def __add__(self, other):
        other = other if isinstance(other, Value) else Value(other)
        out = Value(self.data + other.data, (self, other), "+")
        def _backward():
            self.grad += out.grad                     # += , not =
            other.grad += out.grad
        out._backward = _backward
        return out

    def __mul__(self, other):
        other = other if isinstance(other, Value) else Value(other)
        out = Value(self.data * other.data, (self, other), "*")
        def _backward():
            self.grad += other.data * out.grad
            other.grad += self.data * out.grad
        out._backward = _backward
        return out

    def tanh(self):
        t = math.tanh(self.data)
        out = Value(t, (self,), "tanh")
        def _backward():
            self.grad += (1 - t ** 2) * out.grad
        out._backward = _backward
        return out

    def backward(self):
        order, visited = [], set()
        def build(v):
            if v not in visited:
                visited.add(v)
                for b in v._prev:
                    build(b)
                order.append(v)
        build(self)
        self.grad = 1.0
        for v in reversed(order):                     # topological order, backwards
            v._backward()

    __radd__ = __add__; __rmul__ = __mul__
    def __neg__(self): return self * -1
    def __sub__(self, o): return self + (-o)

# A check: f = (a*b + a).tanh(),  a=2, b=-3  →  f = tanh(-4)
a, b = Value(2.0), Value(-3.0)
f = (a * b + a).tanh()
f.backward()
print(round(f.data, 4), round(a.grad, 4), round(b.grad, 4))
# -0.9993  -0.0013  -0.0007

Always verify numerically: change a.data by ±1e-6, recompute f, and compare the difference quotient with a.grad. If they agree the engine is correct.

Mastery means

  • Implements a scalar autograd engine
  • Builds the computation graph and the topological order
  • Trains a small network with their own engine

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences