Skip to content
AI-grafen
DAI developerClassical machine learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Training, validation and test

Be able to split data correctly, explain why the test data is used only once, and avoid leakage.

Prerequisites

Intuition

Three piles:

  • Training — the model learns here.
  • Validation — you choose the model and the hyperparameters here (compare variants, stop early).
  • Test — used once, at the very end, to report. Never for choosing.

Why? Every time you look at a number and change something, you have «trained» a little on that data — through your choices. Look at the test ten times and the test number is no longer honest.

Leakage = information from the test set or from the future creeping into the training: the same person in both piles, normalisation computed on all the data, features that are only known afterwards.

Code

from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

X_tmp, X_te, y_tmp, y_te = train_test_split(X, y, test_size=0.15, random_state=0, stratify=y)
X_tr, X_va, y_tr, y_va = train_test_split(X_tmp, y_tmp, test_size=0.18, random_state=0, stratify=y_tmp)

sc = StandardScaler().fit(X_tr)          # on the training data only!
X_tr, X_va, X_te = sc.transform(X_tr), sc.transform(X_va), sc.transform(X_te)

# choose the model and hyperparameters with X_va … and RIGHT AT THE END:
print("test:", model.score(X_te, y_te))

stratify preserves the class distribution. For time series: split in time (train on the past, test on the later data), never at random. With several rows per person: split by person (GroupShuffleSplit).

Mastery means

  • Splits the data into training, validation and test sets
  • Explains why the test data is used only once
  • Identifies data leakage

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences