Skip to content
AI-grafen
EUniversityMathematics· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

The chain rule in several variables

Be able to apply the chain rule to composite functions of vectors — what backpropagation is built on.

Prerequisites

Intuition

In one variable: if y = f(g(x)) then dy/dx = f′(g(x)) · g′(x). The derivatives are multiplied along the chain.

In several variables the same holds, but with matrices. A network is a chain:

x → z₁ = W₁x → a₁ = σ(z₁) → z₂ = W₂a₁ → … → L

To know how L changes when W₁ changes, all the intermediate derivatives are multiplied together, from the back. That is backpropagation — not a separate algorithm, but the chain rule applied efficiently.

Derivation

For f:Rn→Rmf:\mathbb R^n\to\mathbb R^m the Jacobian JfJ_f is an m×nm\times n matrix with [Jf]ij=∂fi/∂xj[J_f]_{ij} = \partial f_i/\partial x_j.

The chain rule: for h=f∘gh = f\circ g we have Jh(x)=Jf(g(x)) Jg(x)J_h(x) = J_f(g(x))\, J_g(x) — a matrix product, in that order.

For a scalar loss LL you work with gradients (rows) rather than full Jacobians. For a layer z=Wa+bz = Wa + b, a′=σ(z)a' = \sigma(z):

∂L∂z=∂L∂a′⊙σ′(z),∂L∂W=∂L∂z a⊤,∂L∂a=W⊤∂L∂z\frac{\partial L}{\partial z} = \frac{\partial L}{\partial a'}\odot \sigma'(z), \qquad \frac{\partial L}{\partial W} = \frac{\partial L}{\partial z}\,a^\top, \qquad \frac{\partial L}{\partial a} = W^\top\frac{\partial L}{\partial z}

Why from the back? With nn inputs and a scalar output, forward differentiation (one direction at a time) costs nn passes; backward differentiation gives all the partial derivatives in one pass. With millions of parameters the difference is decisive — it is the whole reason reverse-mode AD dominates in deep learning.

The consequence for deep networks: the gradient reaching layer 1 is a product of many factors. If they are systematically < 1 it vanishes; if they are > 1 it explodes. Residual connections give a «shortcut» where the factor is 1, which is why they make very deep networks trainable.

Code

import numpy as np

def forward(x, W1, b1, W2, b2):
    z1 = W1 @ x + b1
    a1 = np.tanh(z1)
    z2 = W2 @ a1 + b2
    return z1, a1, z2

def backward(x, y, W1, b1, W2, b2):
    z1, a1, z2 = forward(x, W1, b1, W2, b2)
    dL_dz2 = 2 * (z2 - y)                       # L = ||z2 - y||²
    dL_dW2 = np.outer(dL_dz2, a1)
    dL_da1 = W2.T @ dL_dz2                      # the chain rule, backwards
    dL_dz1 = dL_da1 * (1 - np.tanh(z1) ** 2)    # tanh' = 1 - tanh²
    dL_dW1 = np.outer(dL_dz1, x)
    return dL_dW1, dL_dz1, dL_dW2, dL_dz2

# A numerical check — always do this when you have implemented backprop yourself
rng = np.random.default_rng(0)
x, y = rng.normal(size=4), rng.normal(size=2)
W1, b1 = rng.normal(size=(3, 4)) * 0.3, np.zeros(3)
W2, b2 = rng.normal(size=(2, 3)) * 0.3, np.zeros(2)
dW1 = backward(x, y, W1, b1, W2, b2)[0]

eps = 1e-6; num = np.zeros_like(W1)
for i in range(W1.shape[0]):
    for j in range(W1.shape[1]):
        Wp = W1.copy(); Wp[i, j] += eps
        Wm = W1.copy(); Wm[i, j] -= eps
        Lp = ((forward(x, Wp, b1, W2, b2)[2] - y) ** 2).sum()
        Lm = ((forward(x, Wm, b1, W2, b2)[2] - y) ** 2).sum()
        num[i, j] = (Lp - Lm) / (2 * eps)
print(np.abs(dW1 - num).max() < 1e-6)     # True

Mastery means

  • Applies the chain rule to composite vector functions
  • Writes the gradient as a product of Jacobians
  • Connects it to backpropagation

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences