Skip to content
AI-grafen
DAI developerClassical machine learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Overfitting and generalisation

Be able to recognise overfitting in curves and numbers, explain bias/variance, and choose countermeasures.

Prerequisites

Intuition

A model that gets 99 % on the training data and 60 % on new data has memorised instead of learning the pattern. That is overfitting.

The opposite, underfitting: the model is too simple for the pattern — bad on both training and test.

Bias/variance: a simple model → high bias (systematic error), low variance (stable). A complex model → low bias, high variance (the answer swings with whichever examples it happened to see). You want something in between.

The only thing that counts is performance on data the model has not seen.

Code

import numpy as np
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(0)
X = np.sort(rng.uniform(-3, 3, 40))[:, None]
y = np.sin(X[:, 0]) + rng.normal(0, 0.3, 40)
Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.5, random_state=0)

for degree in (1, 3, 15):
    m = make_pipeline(PolynomialFeatures(degree), LinearRegression()).fit(Xtr, ytr)
    print(degree, round(m.score(Xtr, ytr), 2), round(m.score(Xva, yva), 2))
# 1   0.45  0.40   ← underfitted
# 3   0.85  0.80   ← about right
# 15  0.99  -2.1   ← overfitted: perfect on training, a disaster on validation

Formal

The expected test error at a point can be written E[(y−f^(x))2]=Bias[f^(x)]2+Var[f^(x)]+σ2\mathbb{E}[(y - \hat f(x))^2] = \text{Bias}[\hat f(x)]^2 + \text{Var}[\hat f(x)] + \sigma^2, where σ2\sigma^2 is irreducible noise. Model capacity reduces bias but increases variance. Regularisation (L2, say: L+λ∥w∥2L + \lambda\|w\|^2) trades a little bias for much less variance. Early stopping in iterative training is an implicit regularisation.

Mastery means

  • Recognises overfitting in training and validation curves
  • Explains the bias/variance trade-off
  • Chooses a countermeasure: more data, regularisation, a simpler model, early stopping

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences