Overfitting and generalisation
Be able to recognise overfitting in curves and numbers, explain bias/variance, and choose countermeasures.
Prerequisites
Intuition
A model that gets 99 % on the training data and 60 % on new data has memorised instead of learning the pattern. That is overfitting.
The opposite, underfitting: the model is too simple for the pattern — bad on both training and test.
Bias/variance: a simple model → high bias (systematic error), low variance (stable). A complex model → low bias, high variance (the answer swings with whichever examples it happened to see). You want something in between.
The only thing that counts is performance on data the model has not seen.
Code
import numpy as np
from sklearn.preprocessing import PolynomialFeatures
from sklearn.linear_model import LinearRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split
rng = np.random.default_rng(0)
X = np.sort(rng.uniform(-3, 3, 40))[:, None]
y = np.sin(X[:, 0]) + rng.normal(0, 0.3, 40)
Xtr, Xva, ytr, yva = train_test_split(X, y, test_size=0.5, random_state=0)
for degree in (1, 3, 15):
m = make_pipeline(PolynomialFeatures(degree), LinearRegression()).fit(Xtr, ytr)
print(degree, round(m.score(Xtr, ytr), 2), round(m.score(Xva, yva), 2))
# 1 0.45 0.40 ← underfitted
# 3 0.85 0.80 ← about right
# 15 0.99 -2.1 ← overfitted: perfect on training, a disaster on validation
Formal
The expected test error at a point can be written , where is irreducible noise. Model capacity reduces bias but increases variance. Regularisation (L2, say: ) trades a little bias for much less variance. Early stopping in iterative training is an implicit regularisation.
Mastery means
- Recognises overfitting in training and validation curves
- Explains the bias/variance trade-off
- Chooses a countermeasure: more data, regularisation, a simpler model, early stopping
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
Leads to
Part of the goals (24)
- Classical machine learning in practice
- Training neural networks for real
- Seeing and hearing with AI
- Data: collect, clean, document
- Image classification with convolutional networks
- Classical ML for real
- Frontier Lab — an independent research project
- Fine-tune a model with LoRA
- Build a RAG system you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- Fine-tune and run your own models
- Build a memory system for an agent
- Build an agent you can trust
- AI safety in practice
- Evals in practice
- Responsible AI in practice
- Build an AI service that survives production
- An AI service in operation
- Reproduce a paper
- Deep reinforcement learning
- Interpreting a language model
- Statistics for experiments