Multiple comparisons
Be able to correct for many simultaneous tests and explain why benchmark chasing gives false progress.
Prerequisites
- EHypothesis testing and p-valuesrequired
Intuition
Test one hypothesis at α = 0.05 and the risk of a false alarm is 5 %. Test twenty independent hypotheses and the risk that at least one gives a false alarm is:
1 − 0.95²⁰ ≈ 64 %
That is the mathematics behind a large part of what looks like progress in ML. Try twenty architecture variants, report the one that won — and you have reported noise.
| The number of tests | P(at least one false alarm) |
|---|---|
| 1 | 5 % |
| 5 | 23 % |
| 20 | 64 % |
| 100 | 99.4 % |
Formal
Bonferroni: use the threshold for tests. It controls the familywise error rate (the risk of any false alarm). Simple and conservative — with 20 tests the threshold becomes 0.0025, and real effects are easily missed.
Benjamini–Hochberg (FDR): sort the p-values , find the largest with , reject everything up to . It controls the share of false ones among those rejected instead of the risk of any single one. Far more powerful when many tests are done, and the standard in genetics, for instance.
The most important thing is still not the correction but the reporting: report how many variants were tried. A result from «we tested three configurations» and one from «we tested a hundred» are of different strength, even at the same p-value.
Benchmark chasing is the same phenomenon at the field level: hundreds of research groups test against the same benchmark, and those who happen to get high numbers publish. That is why replication on new data is required — not just a record.
Code
import numpy as np
def bonferroni(p, alpha=0.05):
p = np.asarray(p)
return p <= alpha / len(p)
def benjamini_hochberg(p, alpha=0.05):
p = np.asarray(p); m = len(p)
order = np.argsort(p)
thresholds = (np.arange(1, m + 1) / m) * alpha
below = p[order] <= thresholds
k = np.max(np.where(below)[0]) + 1 if below.any() else 0
out = np.zeros(m, bool); out[order[:k]] = True
return out
p = [0.001, 0.008, 0.02, 0.04, 0.06, 0.20, 0.31, 0.55]
print(bonferroni(p).astype(int)) # [1 0 0 0 0 0 0 0]
print(benjamini_hochberg(p).astype(int)) # [1 1 1 0 0 0 0 0]
# the risk of at least one false alarm without a correction
for m in (1, 5, 20, 100):
print(m, round(1 - 0.95 ** m, 3))
Mastery means
- Corrects for many simultaneous tests
- Explains why benchmark chasing gives false progress
- Chooses between Bonferroni and FDR
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Multiple comparisons problem (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — False discovery rate (CC BY-SA 4.0) — CC BY-SA 4.0