Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Hyperparameters for deep networks

Be able to prioritise which hyperparameters matter and set a sweep up.

Prerequisites

Intuition

There are dozens of hyperparameters. Most of them hardly matter, and searching over all of them is a waste.

The priority order, with a rough effect size:

The priorityThe parameterThe effect
1The learning rateenormous — it can decide whether the model learns at all
2The batch size (and the lr along with it)large
3Weight decaymoderate
4The learning rate schedulemoderate
5The model size and depthlarge, but expensive to search
6Dropout and augmentationmoderate, largest with little data
7The optimiser's β valuessmall — leave them alone
8The initialisationsmall with modern defaults

The rule: search over 1–4 and leave the rest at their default values. Searching broadly over everything gives worse results on the same budget than searching carefully over what matters.

And remember: better data nearly always beats better hyperparameters. An hour on data quality more often gives more than a day on sweeps.

Formal

Random search beats grid search. Bergstra & Bengio (2012) showed why: with a grid over two parameters where only one of them matters, you try the same few values of the important parameter over and over. Random search tries a new value every time.

A 3×3 grid (9 runs)              Random (9 runs)
important param: 3 unique values important param: 9 unique values

Search on the right scale. The learning rate and the weight decay should be sampled log-uniformly, not uniformly: the difference between 1e-5 and 1e-4 is as significant as that between 1e-3 and 1e-2.

The parameterThe scaleA typical range
The learning ratelog1e-5 – 1e-2
Weight decaylog1e-6 – 1e-1
Dropoutlinear0 – 0.5
The batch sizelog₂16 – 512
The number of layersinteger2 – 12

The batch size and the learning rate hang together. If you increase the batch you usually need to raise the lr. Two rules of thumb circulate: linear scaling (η∝B\eta \propto B, good for SGD on image networks) and square-root scaling (η∝B\eta \propto \sqrt{B}, often better for Adam). Treat them as starting points, not laws — but do not search the two independently of each other, that wastes budget.

Better than pure random search when the budget is tight:

The methodThe idea
Successive halving / Hyperbandstart many, kill the worst early, give the rest more budget
Bayesian optimisationmodel the result as a function of the parameters and sample where it looks promising
Population based traininglet runs copy and mutate each other's parameters as they go

Hyperband is usually the best trade-off between simplicity and effect.

The trap: overfitting the validation set. If you try 500 configurations and pick the best on the validation set you have in practice trained on it. The difference between the best and the fifth best is then often pure chance. The countermeasure: keep a separate test set that is used only once, and report the spread over a few seeds — not just the best figure.

Code

import numpy as np, itertools, json
from pathlib import Path

SPACE = {
    "lr":           ("log", 1e-5, 1e-2),
    "weight_decay": ("log", 1e-6, 1e-1),
    "dropout":      ("lin", 0.0, 0.5),
    "batch":        ("log2", 4, 9),          # 16 … 512
}

def sample(rng):
    out = {}
    for name, (scale, lo, hi) in SPACE.items():
        if scale == "log":
            out[name] = float(10 ** rng.uniform(np.log10(lo), np.log10(hi)))
        elif scale == "log2":
            out[name] = int(2 ** rng.integers(lo, hi + 1))
        else:
            out[name] = float(rng.uniform(lo, hi))
    return out

def hyperband(train, budget_min=1, budget_max=27, eta=3, seed=0):
    """Start many short runs, keep the best, give them more budget."""
    rng = np.random.default_rng(seed)
    n = int(eta ** np.floor(np.log(budget_max / budget_min) / np.log(eta)))
    candidates = [sample(rng) for _ in range(n)]
    budget = budget_min
    while len(candidates) > 1:
        results = [(train(c, epochs=budget), c) for c in candidates]
        results.sort(key=lambda r: r[0])                  # a lower loss is better
        keep = max(1, len(candidates) // eta)
        print(f"  budget {budget:>3} epochs: {len(candidates)} → {keep} candidates, "
              f"best {results[0][0]:.4f}")
        candidates = [c for _, c in results[:keep]]
        budget *= eta
    return candidates[0]

# Report the spread, not just the best figure
def with_spread(train, params, seeds=(0, 1, 2)):
    v = [train(params, seed=s) for s in seeds]
    return {"mean": round(float(np.mean(v)), 4), "std": round(float(np.std(v)), 4),
            "all": [round(float(x), 4) for x in v]}

# Log everything — otherwise the search cannot be reused
def log(run, params, result, file="sweep.jsonl"):
    with Path(file).open("a", encoding="utf-8") as f:
        f.write(json.dumps({"run": run, **params, **result}, ensure_ascii=False) + "\n")

with_spread is what separates a useful sweep from a misleading one. If the standard deviation between seeds is 0.01 and the difference between the best and the fifth best configuration is 0.005, you have not found a better configuration — you have found a lucky seed.

Mastery means

  • Prioritises hyperparameters by their expected effect
  • Sets a sweep up with the right search type and scale
  • Avoids overfitting the validation set

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences