Hyperparameters for deep networks
Be able to prioritise which hyperparameters matter and set a sweep up.
Prerequisites
- EHyperparameter searchrequired
- ETraining diagnosticsrequired
Intuition
There are dozens of hyperparameters. Most of them hardly matter, and searching over all of them is a waste.
The priority order, with a rough effect size:
| The priority | The parameter | The effect |
|---|---|---|
| 1 | The learning rate | enormous — it can decide whether the model learns at all |
| 2 | The batch size (and the lr along with it) | large |
| 3 | Weight decay | moderate |
| 4 | The learning rate schedule | moderate |
| 5 | The model size and depth | large, but expensive to search |
| 6 | Dropout and augmentation | moderate, largest with little data |
| 7 | The optimiser's β values | small — leave them alone |
| 8 | The initialisation | small with modern defaults |
The rule: search over 1–4 and leave the rest at their default values. Searching broadly over everything gives worse results on the same budget than searching carefully over what matters.
And remember: better data nearly always beats better hyperparameters. An hour on data quality more often gives more than a day on sweeps.
Formal
Random search beats grid search. Bergstra & Bengio (2012) showed why: with a grid over two parameters where only one of them matters, you try the same few values of the important parameter over and over. Random search tries a new value every time.
A 3×3 grid (9 runs) Random (9 runs)
important param: 3 unique values important param: 9 unique values
Search on the right scale. The learning rate and the weight decay should be sampled log-uniformly, not uniformly: the difference between 1e-5 and 1e-4 is as significant as that between 1e-3 and 1e-2.
| The parameter | The scale | A typical range |
|---|---|---|
| The learning rate | log | 1e-5 – 1e-2 |
| Weight decay | log | 1e-6 – 1e-1 |
| Dropout | linear | 0 – 0.5 |
| The batch size | log₂ | 16 – 512 |
| The number of layers | integer | 2 – 12 |
The batch size and the learning rate hang together. If you increase the batch you usually need to raise the lr. Two rules of thumb circulate: linear scaling (, good for SGD on image networks) and square-root scaling (, often better for Adam). Treat them as starting points, not laws — but do not search the two independently of each other, that wastes budget.
Better than pure random search when the budget is tight:
| The method | The idea |
|---|---|
| Successive halving / Hyperband | start many, kill the worst early, give the rest more budget |
| Bayesian optimisation | model the result as a function of the parameters and sample where it looks promising |
| Population based training | let runs copy and mutate each other's parameters as they go |
Hyperband is usually the best trade-off between simplicity and effect.
The trap: overfitting the validation set. If you try 500 configurations and pick the best on the validation set you have in practice trained on it. The difference between the best and the fifth best is then often pure chance. The countermeasure: keep a separate test set that is used only once, and report the spread over a few seeds — not just the best figure.
Code
import numpy as np, itertools, json
from pathlib import Path
SPACE = {
"lr": ("log", 1e-5, 1e-2),
"weight_decay": ("log", 1e-6, 1e-1),
"dropout": ("lin", 0.0, 0.5),
"batch": ("log2", 4, 9), # 16 … 512
}
def sample(rng):
out = {}
for name, (scale, lo, hi) in SPACE.items():
if scale == "log":
out[name] = float(10 ** rng.uniform(np.log10(lo), np.log10(hi)))
elif scale == "log2":
out[name] = int(2 ** rng.integers(lo, hi + 1))
else:
out[name] = float(rng.uniform(lo, hi))
return out
def hyperband(train, budget_min=1, budget_max=27, eta=3, seed=0):
"""Start many short runs, keep the best, give them more budget."""
rng = np.random.default_rng(seed)
n = int(eta ** np.floor(np.log(budget_max / budget_min) / np.log(eta)))
candidates = [sample(rng) for _ in range(n)]
budget = budget_min
while len(candidates) > 1:
results = [(train(c, epochs=budget), c) for c in candidates]
results.sort(key=lambda r: r[0]) # a lower loss is better
keep = max(1, len(candidates) // eta)
print(f" budget {budget:>3} epochs: {len(candidates)} → {keep} candidates, "
f"best {results[0][0]:.4f}")
candidates = [c for _, c in results[:keep]]
budget *= eta
return candidates[0]
# Report the spread, not just the best figure
def with_spread(train, params, seeds=(0, 1, 2)):
v = [train(params, seed=s) for s in seeds]
return {"mean": round(float(np.mean(v)), 4), "std": round(float(np.std(v)), 4),
"all": [round(float(x), 4) for x in v]}
# Log everything — otherwise the search cannot be reused
def log(run, params, result, file="sweep.jsonl"):
with Path(file).open("a", encoding="utf-8") as f:
f.write(json.dumps({"run": run, **params, **result}, ensure_ascii=False) + "\n")
with_spread is what separates a useful sweep from a misleading one. If the standard deviation between seeds is 0.01 and the difference between the best and the fifth best configuration is 0.005, you have not found a better configuration — you have found a lucky seed.
Mastery means
- Prioritises hyperparameters by their expected effect
- Sets a sweep up with the right search type and scale
- Avoids overfitting the validation set
Sign in to do the exercises and build your mastery up.
Sources
- Bergstra & Bengio — Random Search for Hyper-Parameter Optimization (JMLR 2012) — open access
- arXiv — Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization — arXiv (open access; licence per article)
- Google — Deep Learning Tuning Playbook — CC BY 4.0