Jacobians and Hessians
Be able to set up Jacobian and Hessian matrices and understand their role in optimisation.
Prerequisites
Intuition
The gradient generalises the derivative to several variables in, one number out.
The Jacobian generalises it to several in and several out: a matrix where row is the gradient of output .
For , is an matrix. If the Jacobian is just the gradient, lying down.
The Hessian is the second derivatives of a scalar function:
The gradient says which way it slopes. The Hessian says how the surface curves — is it bowl-shaped, dome-shaped or saddle-shaped?
| Matrix | Size | Says |
|---|---|---|
| Gradient | the direction | |
| Jacobian | how each output changes with each input | |
| Hessian | the curvature |
Formal
The Hessian's eigenvalues decide what the point is (at a point where the gradient is zero):
| All the eigenvalues | The point is |
|---|---|
| positive | a local minimum (a bowl) |
| negative | a local maximum (a dome) |
| mixed signs | a saddle point |
| some zero | indeterminate — it requires a higher order |
In high-dimensional problems saddle points are the normal case. With parameters, all eigenvalues have to have the same sign for a genuine local minimum. The probability of that falls fast with , so almost every critical point in a neural network is a saddle. That is good news: gradient descent rarely gets stuck permanently in them, since there is always a direction downwards.
The condition number measures how elongated the bowl is. A large means a narrow ravine — gradient descent zigzags across the edges instead of going forwards, and the learning rate is limited by the steepest direction. That is why normalisation (BatchNorm, LayerNorm) helps: it makes the bowl rounder.
Newton's method uses the curvature to take perfect steps:
It converges quadratically — far faster than gradient descent. So why is it not used in deep learning?
| The obstacle | The size at |
|---|---|
| Storing | numbers — impossible |
| Inverting | — impossible |
| has to be positive definite | it usually is not |
What is actually used are approximations that only need matrix–vector products: L-BFGS (which stores a few vectors), and in practice Adam, which is a crude diagonal approximation of the curvature. Adam's v term estimates the square of the gradients per parameter, which in practice scales each direction by its own curvature — without ever forming a matrix.
Hessian–vector products can be computed without forming , with two backward passes. That is used in pruning, in influence functions and in some interpretability research.
Code
import numpy as np, torch
# The Jacobian of f(x, y) = (x²y, 5x + sin y)
x = torch.tensor([2.0, 1.0], requires_grad=True)
def f(v):
return torch.stack([v[0]**2 * v[1], 5 * v[0] + torch.sin(v[1])])
J = torch.autograd.functional.jacobian(f, x)
print(J)
# tensor([[4.0000, 4.0000], ∂f₁/∂x = 2xy = 4, ∂f₁/∂y = x² = 4
# [5.0000, 0.5403]]) ∂f₂/∂x = 5, ∂f₂/∂y = cos(1)
# The Hessian of a scalar function
def g(v):
return v[0]**2 + 3 * v[0] * v[1] + 2 * v[1]**2
H = torch.autograd.functional.hessian(g, torch.tensor([1.0, 1.0]))
print(H) # [[2., 3.], [3., 4.]]
ev = np.linalg.eigvalsh(H.numpy())
print(np.round(ev, 3)) # [-0.162 6.162] ← mixed signs
print("a saddle point" if ev[0] * ev[-1] < 0 else "a minimum or a maximum")
# The condition number: how elongated is the bowl?
for A in ([[1, 0], [0, 1]], [[1, 0], [0, 100]]):
e = np.linalg.eigvalsh(A)
print(f" eigenvalues {e} condition number {e.max() / e.min():.0f}")
# eigenvalues [1. 1.] condition number 1 ← a round bowl, easy to optimise
# eigenvalues [1. 100.] condition number 100 ← a narrow ravine, zigzagging
# A Hessian–vector product without forming H — O(n) instead of O(n²)
def hvp(f, x, v):
x = x.detach().requires_grad_(True)
grad = torch.autograd.grad(f(x), x, create_graph=True)[0]
return torch.autograd.grad(grad @ v, x)[0]
print(hvp(g, torch.tensor([1.0, 1.0]), torch.tensor([1.0, 0.0]))) # tensor([2., 3.])
# ← the first column of H, without H ever having existed
Mastery means
- Sets up a Jacobian matrix
- Interprets the Hessian matrix and its eigenvalues
- Explains why second-order methods are rarely used in deep learning
Sign in to do the exercises and build your mastery up.
Sources
- Mathematics for Machine Learning (Deisenroth m.fl.) — free to read online (authors' edition)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- PyTorch — autograd — BSD-3-Clause