Data leakage
Be able to detect leakage between training and test and explain why it gives falsely good results.
Prerequisites
- DTraining, validation and testrequired
Intuition
Data leakage = information that is not available at the moment of the decision creeping into the training. The result is brilliant in the evaluation and poor in reality.
Five common forms:
| Type | Example |
|---|---|
| Target leakage | a feature computed from the answer («number_of_reminders» when predicting non-payment) |
| Time leakage | training on the future, testing on the past |
| Group leakage | the same patient or customer in both training and test |
| Preprocessing leakage | scaling or feature selection computed on all the data before the split |
| Duplicates | the same row exists in both sets |
The warning sign: the result is surprisingly good. 0.99 AUC on a hard problem is nearly always leakage, not genius.
Code
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import SelectKBest
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score, GroupKFold
# WRONG: selection and scaling on all the data → leakage
# X_sel = SelectKBest(k=20).fit_transform(X, y)
# cross_val_score(LogisticRegression(), X_sel, y, cv=5) # optimistic
# RIGHT: everything inside the pipeline, fitted per fold
pipe = make_pipeline(StandardScaler(), SelectKBest(k=20), LogisticRegression(max_iter=1000))
print(cross_val_score(pipe, X, y, cv=GroupKFold(5), groups=patient_id).mean())
A checklist before you believe a good result:
- Can every feature really be observed before the thing you are predicting?
- Are there duplicates between training and test? (
set(hash(row))) - Is the split made per group and per time where that is needed?
- Is all the preprocessing inside the pipeline?
- Which single feature gives nearly the whole result? Remove it and see — it is often the leak.
Mastery means
- Gives three examples of data leakage
- Detects leakage when the result is suspiciously good
- Builds pipelines that do not leak
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- arXiv — Leakage and the Reproducibility Crisis in ML-based Science — arXiv (open access; licence per article)