Skip to content
AI-grafen
DAI developerData handling· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Data leakage

Be able to detect leakage between training and test and explain why it gives falsely good results.

Prerequisites

Intuition

Data leakage = information that is not available at the moment of the decision creeping into the training. The result is brilliant in the evaluation and poor in reality.

Five common forms:

TypeExample
Target leakagea feature computed from the answer («number_of_reminders» when predicting non-payment)
Time leakagetraining on the future, testing on the past
Group leakagethe same patient or customer in both training and test
Preprocessing leakagescaling or feature selection computed on all the data before the split
Duplicatesthe same row exists in both sets

The warning sign: the result is surprisingly good. 0.99 AUC on a hard problem is nearly always leakage, not genius.

Code

from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import SelectKBest
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score, GroupKFold

# WRONG: selection and scaling on all the data → leakage
# X_sel = SelectKBest(k=20).fit_transform(X, y)
# cross_val_score(LogisticRegression(), X_sel, y, cv=5)     # optimistic

# RIGHT: everything inside the pipeline, fitted per fold
pipe = make_pipeline(StandardScaler(), SelectKBest(k=20), LogisticRegression(max_iter=1000))
print(cross_val_score(pipe, X, y, cv=GroupKFold(5), groups=patient_id).mean())

A checklist before you believe a good result:

  1. Can every feature really be observed before the thing you are predicting?
  2. Are there duplicates between training and test? (set(hash(row)))
  3. Is the split made per group and per time where that is needed?
  4. Is all the preprocessing inside the pipeline?
  5. Which single feature gives nearly the whole result? Remove it and see — it is often the leak.

Mastery means

  • Gives three examples of data leakage
  • Detects leakage when the result is suspiciously good
  • Builds pipelines that do not leak

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences