Skip to content
AI-grafen
EUniversityStatistics and probability· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

The bootstrap and resampling

Be able to estimate the uncertainty in a model's metric with the bootstrap.

Prerequisites

Intuition

You have one sample and want to know how uncertain your estimate is. Ideally you would draw a hundred new samples from the population — but you only have this one.

The bootstrap: treat your sample as if it were the population. Draw n observations with replacement, compute the metric, repeat a few thousand times. The spread in the thousand values estimates the spread in your estimate.

Why it is so useful in ML: it works for any metric at all — F1, AUC, BLEU, the grounding rate, the cost per solved task, the median latency. Formulas exist only for the simplest metrics; the bootstrap needs none.

Formal

The percentile method (the simplest, usually sufficient): run B ≈ 2 000–10 000 resamplings, sort the estimates, take the 2.5th and the 97.5th percentile.

Variants to know about:

VariantWhen
Percentilethe default choice
BCa (bias-corrected and accelerated)a skewed distribution or bias — more correct coverage
Paired bootstrapcomparing two models on the same cases — resample the cases, keep the pairs
Block bootstraptime series — resample contiguous blocks, not individual points
Stratifiedpreserve the class distribution under imbalance

Where the bootstrap does not work:

  • Extreme values (the maximum, the minimum) — the resampling cannot see beyond the sample's largest value.
  • Very small samples (n < 20) — the resamplings become too much alike.
  • Dependent data without a block structure — independent resampling badly underestimates the uncertainty.

The most common misapplication in ML: resampling the predictions instead of the test cases. The cases are the random unit — and when comparing models the same resampled indices have to be used for both.

Code

import numpy as np
from sklearn.metrics import f1_score

def bootstrap_ci(y, pred, metric=lambda y, p: f1_score(y, p, average="macro"),
                 B=5000, alpha=0.05, seed=0):
    rng = np.random.default_rng(seed)
    n = len(y)
    values = np.empty(B)
    for b in range(B):
        idx = rng.integers(0, n, n)                     # with replacement
        values[b] = metric(y[idx], pred[idx])
    lo, hi = np.percentile(values, [100*alpha/2, 100*(1-alpha/2)])
    return {"estimate": float(metric(y, pred)), "ci": (float(lo), float(hi))}

def paired_bootstrap(y, pred_a, pred_b, metric, B=5000, seed=0):
    """THE SAME indices for both models — otherwise you are measuring the wrong thing."""
    rng = np.random.default_rng(seed); n = len(y)
    diff = np.empty(B)
    for b in range(B):
        idx = rng.integers(0, n, n)
        diff[b] = metric(y[idx], pred_b[idx]) - metric(y[idx], pred_a[idx])
    lo, hi = np.percentile(diff, [2.5, 97.5])
    return {"difference": float(diff.mean()), "ci": (float(lo), float(hi)),
            "share_above_zero": float((diff > 0).mean())}

Mastery means

  • Estimates the uncertainty in a metric with the bootstrap
  • Chooses the right bootstrap variant
  • Knows the method's limits

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences