Learning rate schedules and warm-up
Be able to use cosine, step and warm-up schedules and explain why warm-up helps transformers.
Prerequisites
Intuition
A fixed learning rate is nearly always wrong. At the start small steps are needed (the model is random and the gradients unreliable), in the middle large steps (fast progress), at the end small ones again (fine adjustment).
The standard schedule today: warm-up + a cosine decay.
lr
│ ╭────╮
│ ╱ ╰──╮
│ ╱ ╰───╮
│ ╱ ╰────╮
│ ╱ ╰──────
└─┴────────────────────────────────→ steps
warm-up cosine
(1–5 %)
Why warm-up? Adam's second moment is unreliable in the first steps — it is built on too few observations. A full learning rate then gives large, misdirected steps that can damage the initialisation permanently. Warm-up lets the moment estimates stabilise first.
Code
import math
import numpy as np
def lr_schedule(step, total_steps, max_lr=3e-4, warmup_share=0.03, min_share=0.1):
warmup = int(total_steps * warmup_share)
if step < warmup:
return max_lr * (step + 1) / warmup # a linear warm-up
t = (step - warmup) / max(total_steps - warmup, 1)
cos = 0.5 * (1 + math.cos(math.pi * t)) # 1 → 0
return max_lr * (min_share + (1 - min_share) * cos)
for s in (0, 50, 300, 3000, 7000, 9999):
print(f"step {s:>5}: lr {lr_schedule(s, 10_000):.2e}")
# step 0: lr 1.00e-05
# step 50: lr 1.70e-04
# step 300: lr 3.00e-04 ← the peak, the end of the warm-up
# step 3000: lr 2.49e-04
# step 7000: lr 1.03e-04
# step 9999: lr 3.00e-05 ← min_share · max_lr
Other schedules and when they suit:
| The schedule | Suits |
|---|---|
| Cosine | a known total length — the standard |
| Linear decay | simple, nearly as good |
| Step decay (÷10 at given epochs) | the classic for CNNs |
| OneCycle | fast convergence over few epochs |
| Constant + a decay at the end | an unknown total length (continuous training) |
| ReduceLROnPlateau | when the validation metric is in charge |
The rule: the cosine schedule presupposes that you know total_steps. If you stop early the lr has never come down, and the model has not converged — a common mistake that makes comparisons between different run lengths misleading.
Mastery means
- Uses warm-up and a cosine decay
- Explains why warm-up is needed for transformers
- Chooses the schedule to suit the length of the training
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Attention Is All You Need — arXiv (open access; licence per article)
- arXiv — SGDR: Stochastic Gradient Descent with Warm Restarts — arXiv (open access; licence per article)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0