Skip to content
AI-grafen
EUniversityClassical machine learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

The bias–variance trade-off

Be able to decompose error into bias and variance and connect it to model complexity.

Prerequisites

Intuition

Imagine training the same model type on ten different samples from the same population and looking at the predictions at a given point.

  • Bias = how far the mean of the ten predictions lies from the truth. A systematic error — the model is too simple for the pattern.
  • Variance = how much the ten spread out. The model is so flexible that it follows its particular sample.
  • Noise = what cannot be predicted whatever the model.

A straight line on curved data: high bias, low variance. A degree-15 polynomial: low bias, high variance. Both have a large total error.

Formal

For squared loss at a point xx, with y=f(x)+εy = f(x) + \varepsilon and E[ε]=0\mathbb E[\varepsilon]=0, Var(ε)=σ2\text{Var}(\varepsilon)=\sigma^2:

E[(y−f^(x))2]=(f(x)−E[f^(x)])2⏟bias2+E[(f^(x)−E[f^(x)])2]⏟variance+σ2⏟irreducible\mathbb E\big[(y - \hat f(x))^2\big] = \underbrace{\big(f(x) - \mathbb E[\hat f(x)]\big)^2}_{\text{bias}^2} + \underbrace{\mathbb E\big[(\hat f(x) - \mathbb E[\hat f(x)])^2\big]}_{\text{variance}} + \underbrace{\sigma^2}_{\text{irreducible}}

The expectation is taken over different training sets.

What affects what:

  • Increased model complexity: bias ↓, variance ↑.
  • More training data: variance ↓ (bias unchanged).
  • Regularisation: variance ↓, bias ↑ slightly.
  • Ensembles (bagging, random forest): variance ↓ through averaging.
  • Boosting: bias ↓ by successively correcting errors.

The nuance for modern networks: «double descent» shows that the test error can fall again beyond the interpolation point for very large models — the classic U-curve is not the whole picture.

Code

import numpy as np
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline

rng = np.random.default_rng(0)
f = lambda x: np.sin(1.5 * x)
x_test = np.linspace(-3, 3, 200)[:, None]

for degree in (1, 3, 15):
    preds = []
    for _ in range(50):                        # 50 different training sets
        X = rng.uniform(-3, 3, (30, 1)); y = f(X[:, 0]) + rng.normal(0, 0.3, 30)
        m = make_pipeline(PolynomialFeatures(degree), LinearRegression()).fit(X, y)
        preds.append(m.predict(x_test))
    P = np.array(preds)
    bias2 = float(np.mean((P.mean(axis=0) - f(x_test[:, 0])) ** 2))
    var = float(np.mean(P.var(axis=0)))
    print(f"degree {degree:2d}  bias² {bias2:.3f}  variance {var:.3f}  sum {bias2 + var:.3f}")
# degree  1  bias² 0.281  variance 0.021  sum 0.302
# degree  3  bias² 0.020  variance 0.049  sum 0.069   ← the best
# degree 15  bias² 0.013  variance 0.612  sum 0.625

Mastery means

  • Decomposes error into bias, variance and noise
  • Connects the trade-off to model complexity and the amount of data

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences