Skip to content
AI-grafen
EUniversityMathematics· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Jacobians and Hessians

Be able to set up Jacobian and Hessian matrices and understand their role in optimisation.

Prerequisites

Intuition

The gradient generalises the derivative to several variables in, one number out.

The Jacobian generalises it to several in and several out: a matrix where row ii is the gradient of output ii.

Jij=∂fi∂xjJ_{ij} = \frac{\partial f_i}{\partial x_j}

For f:Rn→Rmf: \mathbb{R}^n \to \mathbb{R}^m, JJ is an m×nm \times n matrix. If m=1m = 1 the Jacobian is just the gradient, lying down.

The Hessian is the second derivatives of a scalar function:

Hij=∂2f∂xi∂xjH_{ij} = \frac{\partial^2 f}{\partial x_i \partial x_j}

The gradient says which way it slopes. The Hessian says how the surface curves — is it bowl-shaped, dome-shaped or saddle-shaped?

MatrixSizeSays
Gradientnnthe direction
Jacobianm×nm \times nhow each output changes with each input
Hessiann×nn \times nthe curvature

Formal

The Hessian's eigenvalues decide what the point is (at a point where the gradient is zero):

All the eigenvaluesThe point is
positivea local minimum (a bowl)
negativea local maximum (a dome)
mixed signsa saddle point
some zeroindeterminate — it requires a higher order

In high-dimensional problems saddle points are the normal case. With nn parameters, all nn eigenvalues have to have the same sign for a genuine local minimum. The probability of that falls fast with nn, so almost every critical point in a neural network is a saddle. That is good news: gradient descent rarely gets stuck permanently in them, since there is always a direction downwards.

The condition number κ=λmax⁡/λmin⁡\kappa = \lambda_{\max}/\lambda_{\min} measures how elongated the bowl is. A large κ\kappa means a narrow ravine — gradient descent zigzags across the edges instead of going forwards, and the learning rate is limited by the steepest direction. That is why normalisation (BatchNorm, LayerNorm) helps: it makes the bowl rounder.

Newton's method uses the curvature to take perfect steps:

xt+1=xt−H−1∇fx_{t+1} = x_t - H^{-1}\nabla f

It converges quadratically — far faster than gradient descent. So why is it not used in deep learning?

The obstacleThe size at n=109n = 10^9
Storing HH101810^{18} numbers — impossible
Inverting HHO(n3)O(n^3) — impossible
HH has to be positive definiteit usually is not

What is actually used are approximations that only need matrix–vector products: L-BFGS (which stores a few vectors), and in practice Adam, which is a crude diagonal approximation of the curvature. Adam's v term estimates the square of the gradients per parameter, which in practice scales each direction by its own curvature — without ever forming a matrix.

Hessian–vector products can be computed without forming HH, with two backward passes. That is used in pruning, in influence functions and in some interpretability research.

Code

import numpy as np, torch

# The Jacobian of f(x, y) = (x²y, 5x + sin y)
x = torch.tensor([2.0, 1.0], requires_grad=True)

def f(v):
    return torch.stack([v[0]**2 * v[1], 5 * v[0] + torch.sin(v[1])])

J = torch.autograd.functional.jacobian(f, x)
print(J)
# tensor([[4.0000, 4.0000],      ∂f₁/∂x = 2xy = 4,  ∂f₁/∂y = x² = 4
#         [5.0000, 0.5403]])     ∂f₂/∂x = 5,        ∂f₂/∂y = cos(1)

# The Hessian of a scalar function
def g(v):
    return v[0]**2 + 3 * v[0] * v[1] + 2 * v[1]**2

H = torch.autograd.functional.hessian(g, torch.tensor([1.0, 1.0]))
print(H)                       # [[2., 3.], [3., 4.]]

ev = np.linalg.eigvalsh(H.numpy())
print(np.round(ev, 3))         # [-0.162  6.162]  ← mixed signs
print("a saddle point" if ev[0] * ev[-1] < 0 else "a minimum or a maximum")

# The condition number: how elongated is the bowl?
for A in ([[1, 0], [0, 1]], [[1, 0], [0, 100]]):
    e = np.linalg.eigvalsh(A)
    print(f"  eigenvalues {e}  condition number {e.max() / e.min():.0f}")
#   eigenvalues [1. 1.]     condition number 1     ← a round bowl, easy to optimise
#   eigenvalues [1. 100.]   condition number 100   ← a narrow ravine, zigzagging

# A Hessian–vector product without forming H — O(n) instead of O(n²)
def hvp(f, x, v):
    x = x.detach().requires_grad_(True)
    grad = torch.autograd.grad(f(x), x, create_graph=True)[0]
    return torch.autograd.grad(grad @ v, x)[0]

print(hvp(g, torch.tensor([1.0, 1.0]), torch.tensor([1.0, 0.0])))   # tensor([2., 3.])
#  ← the first column of H, without H ever having existed

Mastery means

  • Sets up a Jacobian matrix
  • Interprets the Hessian matrix and its eigenvalues
  • Explains why second-order methods are rarely used in deep learning

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences