The chain rule in several variables
Be able to apply the chain rule to composite functions of vectors — what backpropagation is built on.
Prerequisites
Intuition
In one variable: if y = f(g(x)) then dy/dx = f′(g(x)) · g′(x). The derivatives are multiplied along the chain.
In several variables the same holds, but with matrices. A network is a chain:
x → z₁ = W₁x → a₁ = σ(z₁) → z₂ = W₂a₁ → … → L
To know how L changes when W₁ changes, all the intermediate derivatives are multiplied together, from the back. That is backpropagation — not a separate algorithm, but the chain rule applied efficiently.
Derivation
For the Jacobian is an matrix with .
The chain rule: for we have — a matrix product, in that order.
For a scalar loss you work with gradients (rows) rather than full Jacobians. For a layer , :
Why from the back? With inputs and a scalar output, forward differentiation (one direction at a time) costs passes; backward differentiation gives all the partial derivatives in one pass. With millions of parameters the difference is decisive — it is the whole reason reverse-mode AD dominates in deep learning.
The consequence for deep networks: the gradient reaching layer 1 is a product of many factors. If they are systematically < 1 it vanishes; if they are > 1 it explodes. Residual connections give a «shortcut» where the factor is 1, which is why they make very deep networks trainable.
Code
import numpy as np
def forward(x, W1, b1, W2, b2):
z1 = W1 @ x + b1
a1 = np.tanh(z1)
z2 = W2 @ a1 + b2
return z1, a1, z2
def backward(x, y, W1, b1, W2, b2):
z1, a1, z2 = forward(x, W1, b1, W2, b2)
dL_dz2 = 2 * (z2 - y) # L = ||z2 - y||²
dL_dW2 = np.outer(dL_dz2, a1)
dL_da1 = W2.T @ dL_dz2 # the chain rule, backwards
dL_dz1 = dL_da1 * (1 - np.tanh(z1) ** 2) # tanh' = 1 - tanh²
dL_dW1 = np.outer(dL_dz1, x)
return dL_dW1, dL_dz1, dL_dW2, dL_dz2
# A numerical check — always do this when you have implemented backprop yourself
rng = np.random.default_rng(0)
x, y = rng.normal(size=4), rng.normal(size=2)
W1, b1 = rng.normal(size=(3, 4)) * 0.3, np.zeros(3)
W2, b2 = rng.normal(size=(2, 3)) * 0.3, np.zeros(2)
dW1 = backward(x, y, W1, b1, W2, b2)[0]
eps = 1e-6; num = np.zeros_like(W1)
for i in range(W1.shape[0]):
for j in range(W1.shape[1]):
Wp = W1.copy(); Wp[i, j] += eps
Wm = W1.copy(); Wm[i, j] -= eps
Lp = ((forward(x, Wp, b1, W2, b2)[2] - y) ** 2).sum()
Lm = ((forward(x, Wm, b1, W2, b2)[2] - y) ** 2).sum()
num[i, j] = (Lp - Lm) / (2 * eps)
print(np.abs(dW1 - num).max() < 1e-6) # True
Mastery means
- Applies the chain rule to composite vector functions
- Writes the gradient as a product of Jacobians
- Connects it to backpropagation
Sign in to do the exercises and build your mastery up.
Sources
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- PyTorch — Autograd mechanics (BSD-3) — BSD-3-Clause