Better or just different?
Be able to compare two models on the same questions and decide whether the difference is real.
Prerequisites
Intuition
Model A gets 82 % and model B 87 % on your test of 40 questions. Is B better?
Not necessarily. With 40 questions the uncertainty is about ±11 percentage points. A difference of 5 can be pure chance.
But there is a far more sensitive way: compare them pairwise, question by question.
| B right | B wrong | |
|---|---|---|
| A right | 30 | 3 |
| A wrong | 5 | 2 |
The 30 and the 2 where both do the same say nothing about the difference. What counts is 3 against 5 — the cases where they differ. With so few discordant cases the difference is not shown.
Code
import numpy as np
from scipy import stats
a = np.array([1,1,0,1,1,1,0,1,1,1,0,1,1,1,1,0,1,1,1,1,1,0,1,1,1,1,1,0,1,1,1,1,0,1,1,1,1,1,1,0])
b = np.array([1,1,0,1,1,1,1,1,0,1,1,1,1,1,1,0,1,1,1,1,1,1,1,1,1,1,1,0,1,1,1,1,1,1,1,1,1,1,1,0])
print(a.mean(), b.mean()) # 0.80 0.875
b_wins = int(((b == 1) & (a == 0)).sum()) # 5
a_wins = int(((a == 1) & (b == 0)).sum()) # 2
print(a_wins, b_wins, round(stats.binomtest(b_wins, a_wins + b_wins, 0.5).pvalue, 3))
# 2 5 0.453 → not settled
McNemar's test asks: if the models were equally good, how often would 5 against 2 arise by chance? The answer 0.45 means «often» — no conclusion.
How many questions are needed? To detect a difference of 5 percentage points around 85 % you need on the order of 400–600 cases. That is why eval suites are large — and why the same cases are reused between versions.
Mastery means
- Compares two models pairwise on the same questions
- Decides whether the difference is larger than chance
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — McNemar's test (CC BY-SA 4.0) — CC BY-SA 4.0
- arXiv — Show Your Work: Improved Reporting of Experimental Results — arXiv (open access; licence per article)