Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Learning rate schedules and warm-up

Be able to use cosine, step and warm-up schedules and explain why warm-up helps transformers.

Prerequisites

Intuition

A fixed learning rate is nearly always wrong. At the start small steps are needed (the model is random and the gradients unreliable), in the middle large steps (fast progress), at the end small ones again (fine adjustment).

The standard schedule today: warm-up + a cosine decay.

lr
 │      ╭────╮
 │     ╱      ╰──╮
 │    ╱           ╰───╮
 │   ╱                 ╰────╮
 │  ╱                        ╰──────
 └─┴────────────────────────────────→ steps
   warm-up      cosine
   (1–5 %)

Why warm-up? Adam's second moment v^\hat v is unreliable in the first steps — it is built on too few observations. A full learning rate then gives large, misdirected steps that can damage the initialisation permanently. Warm-up lets the moment estimates stabilise first.

Code

import math
import numpy as np

def lr_schedule(step, total_steps, max_lr=3e-4, warmup_share=0.03, min_share=0.1):
    warmup = int(total_steps * warmup_share)
    if step < warmup:
        return max_lr * (step + 1) / warmup                     # a linear warm-up
    t = (step - warmup) / max(total_steps - warmup, 1)
    cos = 0.5 * (1 + math.cos(math.pi * t))                     # 1 → 0
    return max_lr * (min_share + (1 - min_share) * cos)

for s in (0, 50, 300, 3000, 7000, 9999):
    print(f"step {s:>5}: lr {lr_schedule(s, 10_000):.2e}")
# step     0: lr 1.00e-05
# step    50: lr 1.70e-04
# step   300: lr 3.00e-04    ← the peak, the end of the warm-up
# step  3000: lr 2.49e-04
# step  7000: lr 1.03e-04
# step  9999: lr 3.00e-05    ← min_share · max_lr

Other schedules and when they suit:

The scheduleSuits
Cosinea known total length — the standard
Linear decaysimple, nearly as good
Step decay (÷10 at given epochs)the classic for CNNs
OneCyclefast convergence over few epochs
Constant + a decay at the endan unknown total length (continuous training)
ReduceLROnPlateauwhen the validation metric is in charge

The rule: the cosine schedule presupposes that you know total_steps. If you stop early the lr has never come down, and the model has not converged — a common mistake that makes comparisons between different run lengths misleading.

Mastery means

  • Uses warm-up and a cosine decay
  • Explains why warm-up is needed for transformers
  • Chooses the schedule to suit the length of the training

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences