Skip to content
AI-grafen
DAI developerData handling· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Imbalanced classes

Be able to handle skewed classes with reweighting and the right metrics.

Prerequisites

Intuition

1 % of the transactions are fraud. Your model says «no fraud» about everything and gets 99 % accuracy.

It has learnt nothing at all.

Accuracy is useless under imbalance. Use instead:

MetricAnswers
Precisionof the ones we flagged, how many were real?
Recallof the real ones, how many did we find?
F1the harmonic mean of the two
PR-AUCthe whole trade-off, sensitive to the rare class
ROC-AUCthe whole trade-off — but it looks optimistic under strong imbalance

Precision and recall pull in opposite directions. Flag everything and you get perfect recall and terrible precision. Flag only what you are sure of and it is the other way round. Which matters more is decided by what an error costs — not by the statistics.

Formal

Three families of remedy:

RemedyHowComment
Class weightsclass_weight="balanced"the simplest, usually first; changes the loss, not the data
Oversamplingduplicate or generate (SMOTE)a risk of overfitting on the small class
Undersamplingthrow away from the majoritydiscards information
Threshold adjustmentchange 0.5 to something elsefree, and often the most effective

The last row is underrated. A model trained with no remedies at all, but with the threshold set from the cost, often beats a resampled model with a threshold of 0.5.

The threshold should be chosen from the cost, not from symmetry. If a missed fraud costs 5 000 kr and a false alarm costs 50 kr in review time, you should flag as soon as the probability exceeds roughly

p∗=CFPCFP+CFN=5050+5000≈0.01p^* = \frac{C_{FP}}{C_{FP} + C_{FN}} = \frac{50}{50 + 5000} \approx 0.01

That is 1 %, not 50 %. Keeping the default threshold is in practice claiming that the two kinds of error cost the same.

Two rules that are easy to break:

  1. Resample only the training data. Validation and test should have the real class distribution, otherwise you are measuring on a world that does not exist.
  2. Resample inside the cross-validation, not before it. SMOTE before the split creates synthetic points from the test fold and gives leakage.

If the rare class has extremely few examples (tens, not thousands) classification is often the wrong approach. Consider anomaly detection, which models the normal class and flags deviations — it needs no examples of the unusual.

Code

import numpy as np
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (classification_report, average_precision_score,
                             roc_auc_score, precision_recall_curve)

X, y = make_classification(n_samples=20000, weights=[0.99], flip_y=0.01, random_state=0)
Xtr, Xte, ytr, yte = train_test_split(X, y, test_size=0.3, stratify=y, random_state=0)
print("share positive:", round(float(y.mean()), 4))            # 0.0155

for weight in (None, "balanced"):
    m = LogisticRegression(max_iter=2000, class_weight=weight).fit(Xtr, ytr)
    p = m.predict_proba(Xte)[:, 1]
    print(f"class_weight={weight}: ROC-AUC {roc_auc_score(yte, p):.3f}  "
          f"PR-AUC {average_precision_score(yte, p):.3f}")
# class_weight=None:     ROC-AUC 0.962  PR-AUC 0.585
# class_weight=balanced: ROC-AUC 0.962  PR-AUC 0.580
#  ← the weights barely changed the ranking; it is the threshold that does the work

# Choose the threshold from the cost
C_FP, C_FN = 50, 5000
m = LogisticRegression(max_iter=2000).fit(Xtr, ytr)
p = m.predict_proba(Xte)[:, 1]

def cost(t):
    pred = p >= t
    fp = int(((pred == 1) & (yte == 0)).sum())
    fn = int(((pred == 0) & (yte == 1)).sum())
    return fp * C_FP + fn * C_FN, fp, fn

for t in (0.5, 0.2, 0.05, C_FP / (C_FP + C_FN)):
    c, fp, fn = cost(t)
    print(f"  threshold {t:.3f}: cost {c:>8,} kr  (FP {fp:>4}, FN {fn:>3})")

best = min(np.linspace(0.001, 0.999, 999), key=lambda t: cost(t)[0])
print("cheapest threshold:", round(float(best), 3))

The output shows the important part: the difference between a threshold of 0.5 and a cost-chosen threshold is usually larger than the difference between models. And it costs no training time at all.

Mastery means

  • Chooses metrics that work under imbalance
  • Uses class weights or resampling
  • Sets the decision threshold from the cost

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences