Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Optimisers: momentum, Adam, scheduling

Be able to explain momentum and Adam, choose a learning rate and a schedule, and debug a training run that does not converge.

Prerequisites

Intuition

SGD takes a step towards minus the gradient. The problem: the gradient jumps between batches, and in a long narrow valley it bounces between the walls instead of going forward.

Momentum adds inertia: the step is a moving average of the earlier gradients. The bounces cancel each other out, the forward direction is reinforced.

Adam adds one more thing: every parameter gets its own step length, based on how large its gradients usually are. Parameters with small gradients get larger steps. That makes Adam robust against badly scaled problems — which is why it is the default choice.

The learning rate is still the most important thing. Too large: the loss becomes NaN or bounces. Too small: nothing happens. Typical values: 3e-4 for transformers, 1e-3 for small networks, 2e-5 for fine-tuning.

Formal

Momentum: vt=βvt−1+gtv_t = \beta v_{t-1} + g_t, θt=θt−1−ηvt\theta_t = \theta_{t-1} - \eta v_t with β≈0.9\beta \approx 0.9.

Adam: keeps two moving averages — the first moment (the direction) and the second moment (the size):

mt=β1mt−1+(1−β1)gtm_t = \beta_1 m_{t-1} + (1-\beta_1)g_t, vt=β2vt−1+(1−β2)gt2v_t = \beta_2 v_{t-1} + (1-\beta_2)g_t^2

Bias correction (important at the start, when the averages begin at zero): m^t=mt/(1−β1t)\hat m_t = m_t/(1-\beta_1^t), v^t=vt/(1−β2t)\hat v_t = v_t/(1-\beta_2^t).

The update: θt=θt−1−η m^t/(v^t+ϵ)\theta_t = \theta_{t-1} - \eta\,\hat m_t/(\sqrt{\hat v_t}+\epsilon) with β1=0.9\beta_1 = 0.9, β2=0.999\beta_2 = 0.999, ϵ=10−8\epsilon = 10^{-8}.

AdamW separates the weight decay from the gradient (θ←θ−ηλθ\theta \leftarrow \theta - \eta\lambda\theta separately) — that is correct regularisation and the standard for transformers.

The schedule: a linear warm-up over a few hundred steps (otherwise the first Adam steps are unstable) followed by a cosine decay towards zero.

Code

import torch

opt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01, betas=(0.9, 0.999))
sched = torch.optim.lr_scheduler.OneCycleLR(opt, max_lr=3e-4, total_steps=total_steps, pct_start=0.05)

for xb, yb in loader:
    loss = criterion(model(xb), yb)
    loss.backward()
    torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)   # against exploding gradients
    opt.step(); sched.step(); opt.zero_grad()

A debugging scheme for when the training does not converge:

The symptomThe likely causeThe measure
loss = NaN immediatelylr too high, or a division by zero in the datalower lr 10×, check the data
the loss bounceslr too high, the batch too smalllower lr, increase the batch
the loss stands stilllr too low, a dead ReLU, the wrong lossraise lr, check the norm of the gradients
the loss falls but the validation risesoverfittingregularise, early stopping
the loss does not fall even on 10 examplesa bug, not a hyperparameterdeliberately overfit a small batch first

The last row is the best test there is: a correct implementation should be able to get the loss close to zero on ten examples. If it cannot, the fault is in the code.

Mastery means

  • Explains momentum and Adam's two moments
  • Chooses a learning rate and a schedule
  • Debugs a training run that does not converge

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences