Imbalanced classes
Be able to handle skewed classes with reweighting and the right metrics.
Prerequisites
- BClassification: how a model sorts informationrequired
- DFeature engineeringrequired
Intuition
1 % of the transactions are fraud. Your model says «no fraud» about everything and gets 99 % accuracy.
It has learnt nothing at all.
Accuracy is useless under imbalance. Use instead:
| Metric | Answers |
|---|---|
| Precision | of the ones we flagged, how many were real? |
| Recall | of the real ones, how many did we find? |
| F1 | the harmonic mean of the two |
| PR-AUC | the whole trade-off, sensitive to the rare class |
| ROC-AUC | the whole trade-off — but it looks optimistic under strong imbalance |
Precision and recall pull in opposite directions. Flag everything and you get perfect recall and terrible precision. Flag only what you are sure of and it is the other way round. Which matters more is decided by what an error costs — not by the statistics.
Formal
Three families of remedy:
| Remedy | How | Comment |
|---|---|---|
| Class weights | class_weight="balanced" | the simplest, usually first; changes the loss, not the data |
| Oversampling | duplicate or generate (SMOTE) | a risk of overfitting on the small class |
| Undersampling | throw away from the majority | discards information |
| Threshold adjustment | change 0.5 to something else | free, and often the most effective |
The last row is underrated. A model trained with no remedies at all, but with the threshold set from the cost, often beats a resampled model with a threshold of 0.5.
The threshold should be chosen from the cost, not from symmetry. If a missed fraud costs 5 000 kr and a false alarm costs 50 kr in review time, you should flag as soon as the probability exceeds roughly
That is 1 %, not 50 %. Keeping the default threshold is in practice claiming that the two kinds of error cost the same.
Two rules that are easy to break:
- Resample only the training data. Validation and test should have the real class distribution, otherwise you are measuring on a world that does not exist.
- Resample inside the cross-validation, not before it. SMOTE before the split creates synthetic points from the test fold and gives leakage.
If the rare class has extremely few examples (tens, not thousands) classification is often the wrong approach. Consider anomaly detection, which models the normal class and flags deviations — it needs no examples of the unusual.
Code
import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (classification_report, average_precision_score,
roc_auc_score, precision_recall_curve)
X, y = make_classification(n_samples=20000, weights=[0.99], flip_y=0.01, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, stratify=y, random_state=0)
print("share positive:", round(float(y.mean()), 4)) # 0.0155
for weight in (None, "balanced"):
m = LogisticRegression(max_iter=2000, class_weight=weight).fit(Xtr, ytr)
p = m.predict_proba(Xte)[:, 1]
print(f"class_weight={weight}: ROC-AUC {roc_auc_score(yte, p):.3f} "
f"PR-AUC {average_precision_score(yte, p):.3f}")
# class_weight=None: ROC-AUC 0.962 PR-AUC 0.585
# class_weight=balanced: ROC-AUC 0.962 PR-AUC 0.580
# ← the weights barely changed the ranking; it is the threshold that does the work
# Choose the threshold from the cost
C_FP, C_FN = 50, 5000
m = LogisticRegression(max_iter=2000).fit(Xtr, ytr)
p = m.predict_proba(Xte)[:, 1]
def cost(t):
pred = p >= t
fp = int(((pred == 1) & (yte == 0)).sum())
fn = int(((pred == 0) & (yte == 1)).sum())
return fp * C_FP + fn * C_FN, fp, fn
for t in (0.5, 0.2, 0.05, C_FP / (C_FP + C_FN)):
c, fp, fn = cost(t)
print(f" threshold {t:.3f}: cost {c:>8,} kr (FP {fp:>4}, FN {fn:>3})")
best = min(np.linspace(0.001, 0.999, 999), key=lambda t: cost(t)[0])
print("cheapest threshold:", round(float(best), 3))
The output shows the important part: the difference between a threshold of 0.5 and a cost-chosen threshold is usually larger than the difference between models. And it costs no training time at all.
Mastery means
- Chooses metrics that work under imbalance
- Uses class weights or resampling
- Sets the decision threshold from the cost
Sign in to do the exercises and build your mastery up.
Sources
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Google — Rules of Machine Learning — CC BY 4.0
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0