FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Statistical significance in evals
Be able to decide whether a difference between two models is real.
Prerequisites
- EBuild an eval harnessrequired
- EConfidence intervalsrequired
Intuition
Two models, the same 200 eval cases, 84 % against 87 %. A real difference or noise?
Do three things, in this order:
- Paired, not independent. The same cases for both → look at the cases where they differ. With 200 cases and 12 discordant, the total says far less than those 12.
- The bootstrap for an interval. It works for any metric at all — accuracy, F1, the grounding rate, the cost per solved task.
- Report the effect size with an interval, not just p. «+3.0 pp [0.4, 5.6]» says what you need to know.
Code
import numpy as np
from scipy import stats
def paired_bootstrap(a, b, n=10_000, seed=0):
"""a, b: binary outcomes per case (the same order). Returns the difference with a 95 % interval."""
a, b = np.asarray(a, float), np.asarray(b, float)
d = b - a
rng = np.random.default_rng(seed)
idx = rng.integers(0, len(d), size=(n, len(d)))
distribution = d[idx].mean(axis=1)
lo, hi = np.percentile(distribution, [2.5, 97.5])
return {"difference": float(d.mean()), "ci": (float(lo), float(hi)),
"share_above_zero": float((distribution > 0).mean())}
def mcnemar(a, b):
b_only = int(((b == 1) & (a == 0)).sum())
a_only = int(((a == 1) & (b == 0)).sum())
p = stats.binomtest(b_only, b_only + a_only, 0.5).pvalue if (b_only + a_only) else 1.0
return {"b_wins": b_only, "a_wins": a_only, "p": float(p)}
rng = np.random.default_rng(0)
a = rng.binomial(1, 0.84, 200); b = a.copy()
b[rng.choice(np.where(a == 0)[0], 12, replace=False)] = 1
b[rng.choice(np.where(a == 1)[0], 6, replace=False)] = 0
print(paired_bootstrap(a, b)) # {'difference': 0.03, 'ci': (0.0, 0.06), ...}
print(mcnemar(a, b)) # {'b_wins': 12, 'a_wins': 6, 'p': 0.238}
How many cases are needed? To detect a difference of 3 percentage points around 85 % with a paired test you need on the order of 400–600 cases. With 100 you can only detect differences of about 8 percentage points — which is worth knowing before you interpret the result.
Mastery means
- Uses paired tests on eval results
- Computes bootstrap intervals for arbitrary metrics
- Reports the effect size with its uncertainty
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — McNemar's test (CC BY-SA 4.0) — CC BY-SA 4.0
- arXiv — Show Your Work: Improved Reporting of Experimental Results — arXiv (open access; licence per article)