The bootstrap and resampling
Be able to estimate the uncertainty in a model's metric with the bootstrap.
Prerequisites
- EConfidence intervalsrequired
- EMonte Carlo methodsrequired
Intuition
You have one sample and want to know how uncertain your estimate is. Ideally you would draw a hundred new samples from the population — but you only have this one.
The bootstrap: treat your sample as if it were the population. Draw n observations with replacement, compute the metric, repeat a few thousand times. The spread in the thousand values estimates the spread in your estimate.
Why it is so useful in ML: it works for any metric at all — F1, AUC, BLEU, the grounding rate, the cost per solved task, the median latency. Formulas exist only for the simplest metrics; the bootstrap needs none.
Formal
The percentile method (the simplest, usually sufficient): run B ≈ 2 000–10 000 resamplings, sort the estimates, take the 2.5th and the 97.5th percentile.
Variants to know about:
| Variant | When |
|---|---|
| Percentile | the default choice |
| BCa (bias-corrected and accelerated) | a skewed distribution or bias — more correct coverage |
| Paired bootstrap | comparing two models on the same cases — resample the cases, keep the pairs |
| Block bootstrap | time series — resample contiguous blocks, not individual points |
| Stratified | preserve the class distribution under imbalance |
Where the bootstrap does not work:
- Extreme values (the maximum, the minimum) — the resampling cannot see beyond the sample's largest value.
- Very small samples (n < 20) — the resamplings become too much alike.
- Dependent data without a block structure — independent resampling badly underestimates the uncertainty.
The most common misapplication in ML: resampling the predictions instead of the test cases. The cases are the random unit — and when comparing models the same resampled indices have to be used for both.
Code
import numpy as np
from sklearn.metrics import f1_score
def bootstrap_ci(y, pred, metric=lambda y, p: f1_score(y, p, average="macro"),
B=5000, alpha=0.05, seed=0):
rng = np.random.default_rng(seed)
n = len(y)
values = np.empty(B)
for b in range(B):
idx = rng.integers(0, n, n) # with replacement
values[b] = metric(y[idx], pred[idx])
lo, hi = np.percentile(values, [100*alpha/2, 100*(1-alpha/2)])
return {"estimate": float(metric(y, pred)), "ci": (float(lo), float(hi))}
def paired_bootstrap(y, pred_a, pred_b, metric, B=5000, seed=0):
"""THE SAME indices for both models — otherwise you are measuring the wrong thing."""
rng = np.random.default_rng(seed); n = len(y)
diff = np.empty(B)
for b in range(B):
idx = rng.integers(0, n, n)
diff[b] = metric(y[idx], pred_b[idx]) - metric(y[idx], pred_a[idx])
lo, hi = np.percentile(diff, [2.5, 97.5])
return {"difference": float(diff.mean()), "ci": (float(lo), float(hi)),
"share_above_zero": float((diff > 0).mean())}
Mastery means
- Estimates the uncertainty in a metric with the bootstrap
- Chooses the right bootstrap variant
- Knows the method's limits
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Bootstrapping (statistics) (CC BY-SA 4.0) — CC BY-SA 4.0
- Efron & Tibshirani — An Introduction to the Bootstrap (referens) — CC BY-SA 4.0 (uppslagsverk)