Data drift and model decay
Be able to detect drift and decide when a model should be retrained.
Prerequisites
- DProbability distributionsrequired
- EObservability for ML systemsrequired
Intuition
A model trained on yesterday's world meets tomorrow's. Three kinds of change, with different remedies:
| Type | What changes | Example | Remedy |
|---|---|---|---|
| Covariate drift | P(X) | new user groups, new phrasing | retrain on new data |
| Label drift | P(Y) | the share of fraud rises | adjust the threshold or the weights |
| Concept drift | P(Y|X) | the same input means something else | retrain — there is no shortcut |
Concept drift is the worst: the model cannot detect it itself, since the input looks normal. Only the ground truth reveals it — and the ground truth often arrives late.
Code
import numpy as np
from scipy.stats import ks_2samp, chi2_contingency
def psi(reference, new, bins=10):
"""The Population Stability Index. < 0.1 stable · 0.1–0.25 moderate · > 0.25 substantial drift."""
edges = np.percentile(reference, np.linspace(0, 100, bins + 1))
edges[0], edges[-1] = -np.inf, np.inf
r = np.histogram(reference, edges)[0] / len(reference)
n = np.histogram(new, edges)[0] / len(new)
r, n = np.clip(r, 1e-6, None), np.clip(n, 1e-6, None)
return float(((n - r) * np.log(n / r)).sum())
def drift_report(ref_df, new_df, numeric, categorical):
out = []
for k in numeric:
p = ks_2samp(ref_df[k], new_df[k]).pvalue
out.append({"feature": k, "test": "KS", "p": round(float(p), 5),
"psi": round(psi(ref_df[k], new_df[k]), 3)})
for k in categorical:
tab = np.array([ref_df[k].value_counts().sort_index(),
new_df[k].value_counts().sort_index()])
out.append({"feature": k, "test": "chi2", "p": round(float(chi2_contingency(tab).pvalue), 5)})
return sorted(out, key=lambda r: r.get("psi", 0), reverse=True)
Set the threshold on the PSI, not on the p-value. With large enough datasets every test becomes significant; the PSI measures how much the distribution has moved, which is what matters.
When should you retrain? Decide the rule in advance: (1) on a schedule, monthly say, (2) at a PSI above 0.25 on an important feature, or (3) at a measurable quality degradation on a labelled control stream. The third is the best but requires ongoing labels — budget for labelling 1–2 % of the traffic.
Mastery means
- Tells covariate, label and concept drift apart
- Detects drift with suitable tests
- Decides when the model should be retrained
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Concept drift (CC BY-SA 4.0) — CC BY-SA 4.0
- Google — Rules of Machine Learning — CC BY 4.0