Cross-validation
Be able to use k-fold cross-validation and understand when it is needed.
Prerequisites
- DTraining, validation and testrequired
Intuition
With little data a single validation split becomes unreliable: 200 examples split 80/20 gives 40 validation examples, and the result swings several percentage points depending on which 40 ended up there.
k-fold cross-validation: divide the data into k equal parts. Train k times, each time with one part as the validation set and the rest as training. The mean of the k results is much more stable — and the spread says how uncertain the figure is.
k = 5 or 10 is standard. The cost is that you train k times.
Code
from sklearn.model_selection import cross_val_score, StratifiedKFold, GroupKFold
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
# A pipeline inside cross_val_score → the scaling is fitted within each fold. Without a
# pipeline the validation part's mean leaks into the training.
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
s = cross_val_score(model, X, y, cv=cv, scoring="f1_macro")
print(f"{s.mean():.3f} ± {s.std():.3f}") # 0.812 ± 0.024
# Several rows per person? Split per person, or the individual leaks between the folds:
# cross_val_score(model, X, y, cv=GroupKFold(5), groups=person_id)
Nested cross-validation is needed if you both choose hyperparameters and want an honest estimate: an outer loop for the estimate, an inner one for the choice. Otherwise the result is optimistic, since the choice was made on the same data that measures it.
Mastery means
- Uses k-fold cross-validation
- Explains when it is needed and what it costs
- Avoids leakage in the cross-validation
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause