Backpropagation
Be able to derive the gradient for the weights in a small network with the chain rule, explain why you go backwards, and implement backprop for a two-layer network without autograd.
Prerequisites
- DGradient descentrequired
- DNeural networks — the forward pass with matricesrequired
Intuition
We know how to walk down a loss curve if we have the gradient. The problem: a network has millions of weights, and the loss depends on every weight through a long chain of layers.
Backpropagation is the chain rule organised cleverly: work out the loss's sensitivity to the output, send it backwards layer by layer, and multiply by each layer's local slope. Every intermediate result is reused — which is why it is cheap.
Derivation
The network: z₁ = W₁x + b₁, a₁ = ReLU(z₁), z₂ = W₂a₁ + b₂, L = ½‖z₂ − y‖².
Backwards:
- δ₂ = ∂L/∂z₂ = z₂ − y
- ∂L/∂W₂ = δ₂ a₁ᵀ, ∂L/∂b₂ = δ₂
- δ₁ = ∂L/∂z₁ = (W₂ᵀ δ₂) ⊙ ReLU'(z₁), where ReLU'(z) = 1 if z > 0 else 0
- ∂L/∂W₁ = δ₁ xᵀ, ∂L/∂b₁ = δ₁
The pattern: δ for a layer = (the next layer's weights)ᵀ · (the next layer's δ) ⊙ the local derivative. The gradient for a weight matrix = (the layer's δ) · (the layer's input)ᵀ.
Code
import numpy as np
rng = np.random.default_rng(1)
x, y = rng.random(4), rng.random(2)
W1, b1 = rng.normal(size=(3, 4)), np.zeros(3)
W2, b2 = rng.normal(size=(2, 3)), np.zeros(2)
z1 = W1 @ x + b1; a1 = np.maximum(0, z1)
z2 = W2 @ a1 + b2; L = 0.5 * ((z2 - y) ** 2).sum()
d2 = z2 - y
dW2, db2 = np.outer(d2, a1), d2
d1 = (W2.T @ d2) * (z1 > 0)
dW1, db1 = np.outer(d1, x), d1
# a numerical check of one weight
eps = 1e-5; W1c = W1.copy(); W1c[0, 0] += eps
z2c = W2 @ np.maximum(0, W1c @ x + b1) + b2
Lc = 0.5 * ((z2c - y) ** 2).sum()
print((Lc - L) / eps, dW1[0, 0]) # should be almost the same
The numerical check is your friend: if it does not match, you have an error in the derivation.
Mastery means
- Derives ∂L/∂W for the last layer
- Implements the backward pass for a two-layer network and verifies it numerically
Sign in to do the exercises and build your mastery up.
Sources
Leads to
Part of the goals (29)
- Training neural networks for real
- Fine-tune a model with LoRA
- Train your first neural network
- Fine-tune and run your own models
- Frontier Lab — an independent research project
- Multimodal systems
- Classical ML for real
- Build a transformer from scratch
- Understand how generative AI works
- Image classification with convolutional networks
- Build a voice interface
- Run models more cheaply: quantisation
- Language models in practice
- AI safety in practice
- Responsible AI in practice
- Build a RAG system you can trust
- Statistics for experiments
- Reproduce a paper
- Evals in practice
- Interpreting a language model
- Deep reinforcement learning
- Build an agent you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- An AI service in operation
- Build a memory system for an agent
- Build an AI service that survives production
- AI, ethics and society