Build a mini autograd
Be able to implement a scalar autograd engine (micrograd style) and train a network with it.
Prerequisites
- DPython — classes and objectsrequired
- EHow autograd works insiderequired
Intuition
An autograd engine needs surprisingly little:
- A Value class that carries a number, its gradient, its «children» and a local
_backwardfunction. - Every operation (+, ·, tanh …) creates a new Value, connects it to its operands and defines how the gradient is to be propagated backwards through that particular operation.
backward()sorts the graph topologically and runs_backwardin reverse order.
That is all. PyTorch does the same thing, but on tensors and with a thousand optimisations.
The key insight: gradients are accumulated (+=). A variable used twice gets a contribution from both paths — that is the chain rule's sum rule.
Code
import math
class Value:
def __init__(self, data, children=(), op=""):
self.data, self.grad = float(data), 0.0
self._backward = lambda: None
self._prev, self._op = set(children), op
def __add__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data + other.data, (self, other), "+")
def _backward():
self.grad += out.grad # += , not =
other.grad += out.grad
out._backward = _backward
return out
def __mul__(self, other):
other = other if isinstance(other, Value) else Value(other)
out = Value(self.data * other.data, (self, other), "*")
def _backward():
self.grad += other.data * out.grad
other.grad += self.data * out.grad
out._backward = _backward
return out
def tanh(self):
t = math.tanh(self.data)
out = Value(t, (self,), "tanh")
def _backward():
self.grad += (1 - t ** 2) * out.grad
out._backward = _backward
return out
def backward(self):
order, visited = [], set()
def build(v):
if v not in visited:
visited.add(v)
for b in v._prev:
build(b)
order.append(v)
build(self)
self.grad = 1.0
for v in reversed(order): # topological order, backwards
v._backward()
__radd__ = __add__; __rmul__ = __mul__
def __neg__(self): return self * -1
def __sub__(self, o): return self + (-o)
# A check: f = (a*b + a).tanh(), a=2, b=-3 → f = tanh(-4)
a, b = Value(2.0), Value(-3.0)
f = (a * b + a).tanh()
f.backward()
print(round(f.data, 4), round(a.grad, 4), round(b.grad, 4))
# -0.9993 -0.0013 -0.0007
Always verify numerically: change a.data by ±1e-6, recompute f, and compare the difference quotient with a.grad. If they agree the engine is correct.
Mastery means
- Implements a scalar autograd engine
- Builds the computation graph and the topological order
- Trains a small network with their own engine
Sign in to do the exercises and build your mastery up.
Sources
- Karpathy — micrograd (MIT) — MIT
- PyTorch — Autograd mechanics (BSD-3) — BSD-3-Clause