Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Statistical significance in evals

Be able to decide whether a difference between two models is real.

Prerequisites

Intuition

Two models, the same 200 eval cases, 84 % against 87 %. A real difference or noise?

Do three things, in this order:

  1. Paired, not independent. The same cases for both → look at the cases where they differ. With 200 cases and 12 discordant, the total says far less than those 12.
  2. The bootstrap for an interval. It works for any metric at all — accuracy, F1, the grounding rate, the cost per solved task.
  3. Report the effect size with an interval, not just p. «+3.0 pp [0.4, 5.6]» says what you need to know.

Code

import numpy as np
from scipy import stats

def paired_bootstrap(a, b, n=10_000, seed=0):
    """a, b: binary outcomes per case (the same order). Returns the difference with a 95 % interval."""
    a, b = np.asarray(a, float), np.asarray(b, float)
    d = b - a
    rng = np.random.default_rng(seed)
    idx = rng.integers(0, len(d), size=(n, len(d)))
    distribution = d[idx].mean(axis=1)
    lo, hi = np.percentile(distribution, [2.5, 97.5])
    return {"difference": float(d.mean()), "ci": (float(lo), float(hi)),
            "share_above_zero": float((distribution > 0).mean())}

def mcnemar(a, b):
    b_only = int(((b == 1) & (a == 0)).sum())
    a_only = int(((a == 1) & (b == 0)).sum())
    p = stats.binomtest(b_only, b_only + a_only, 0.5).pvalue if (b_only + a_only) else 1.0
    return {"b_wins": b_only, "a_wins": a_only, "p": float(p)}

rng = np.random.default_rng(0)
a = rng.binomial(1, 0.84, 200); b = a.copy()
b[rng.choice(np.where(a == 0)[0], 12, replace=False)] = 1
b[rng.choice(np.where(a == 1)[0], 6, replace=False)] = 0
print(paired_bootstrap(a, b))    # {'difference': 0.03, 'ci': (0.0, 0.06), ...}
print(mcnemar(a, b))             # {'b_wins': 12, 'a_wins': 6, 'p': 0.238}

How many cases are needed? To detect a difference of 3 percentage points around 85 % with a paired test you need on the order of 400–600 cases. With 100 you can only detect differences of about 8 percentage points — which is worth knowing before you interpret the result.

Mastery means

  • Uses paired tests on eval results
  • Computes bootstrap intervals for arbitrary metrics
  • Reports the effect size with its uncertainty

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences