Skip to content
AI-grafen
FAI engineeringStatistics and probability· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Multiple comparisons

Be able to correct for many simultaneous tests and explain why benchmark chasing gives false progress.

Prerequisites

Intuition

Test one hypothesis at α = 0.05 and the risk of a false alarm is 5 %. Test twenty independent hypotheses and the risk that at least one gives a false alarm is:

1 − 0.95²⁰ ≈ 64 %

That is the mathematics behind a large part of what looks like progress in ML. Try twenty architecture variants, report the one that won — and you have reported noise.

The number of testsP(at least one false alarm)
15 %
523 %
2064 %
10099.4 %

Formal

Bonferroni: use the threshold α/m\alpha/m for mm tests. It controls the familywise error rate (the risk of any false alarm). Simple and conservative — with 20 tests the threshold becomes 0.0025, and real effects are easily missed.

Benjamini–Hochberg (FDR): sort the p-values p(1)≤⋯≤p(m)p_{(1)}\le\dots\le p_{(m)}, find the largest kk with p(k)≤kmαp_{(k)} \le \frac{k}{m}\alpha, reject everything up to kk. It controls the share of false ones among those rejected instead of the risk of any single one. Far more powerful when many tests are done, and the standard in genetics, for instance.

The most important thing is still not the correction but the reporting: report how many variants were tried. A result from «we tested three configurations» and one from «we tested a hundred» are of different strength, even at the same p-value.

Benchmark chasing is the same phenomenon at the field level: hundreds of research groups test against the same benchmark, and those who happen to get high numbers publish. That is why replication on new data is required — not just a record.

Code

import numpy as np

def bonferroni(p, alpha=0.05):
    p = np.asarray(p)
    return p <= alpha / len(p)

def benjamini_hochberg(p, alpha=0.05):
    p = np.asarray(p); m = len(p)
    order = np.argsort(p)
    thresholds = (np.arange(1, m + 1) / m) * alpha
    below = p[order] <= thresholds
    k = np.max(np.where(below)[0]) + 1 if below.any() else 0
    out = np.zeros(m, bool); out[order[:k]] = True
    return out

p = [0.001, 0.008, 0.02, 0.04, 0.06, 0.20, 0.31, 0.55]
print(bonferroni(p).astype(int))            # [1 0 0 0 0 0 0 0]
print(benjamini_hochberg(p).astype(int))    # [1 1 1 0 0 0 0 0]

# the risk of at least one false alarm without a correction
for m in (1, 5, 20, 100):
    print(m, round(1 - 0.95 ** m, 3))

Mastery means

  • Corrects for many simultaneous tests
  • Explains why benchmark chasing gives false progress
  • Chooses between Bonferroni and FDR

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences