Skip to content
AI-grafen
DAI developerEvals and benchmarks· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Better or just different?

Be able to compare two models on the same questions and decide whether the difference is real.

Prerequisites

Intuition

Model A gets 82 % and model B 87 % on your test of 40 questions. Is B better?

Not necessarily. With 40 questions the uncertainty is about ±11 percentage points. A difference of 5 can be pure chance.

But there is a far more sensitive way: compare them pairwise, question by question.

B rightB wrong
A right303
A wrong52

The 30 and the 2 where both do the same say nothing about the difference. What counts is 3 against 5 — the cases where they differ. With so few discordant cases the difference is not shown.

Code

import numpy as np
from scipy import stats

a = np.array([1,1,0,1,1,1,0,1,1,1,0,1,1,1,1,0,1,1,1,1,1,0,1,1,1,1,1,0,1,1,1,1,0,1,1,1,1,1,1,0])
b = np.array([1,1,0,1,1,1,1,1,0,1,1,1,1,1,1,0,1,1,1,1,1,1,1,1,1,1,1,0,1,1,1,1,1,1,1,1,1,1,1,0])

print(a.mean(), b.mean())                       # 0.80 0.875
b_wins = int(((b == 1) & (a == 0)).sum())       # 5
a_wins = int(((a == 1) & (b == 0)).sum())       # 2
print(a_wins, b_wins, round(stats.binomtest(b_wins, a_wins + b_wins, 0.5).pvalue, 3))
# 2 5 0.453  → not settled

McNemar's test asks: if the models were equally good, how often would 5 against 2 arise by chance? The answer 0.45 means «often» — no conclusion.

How many questions are needed? To detect a difference of 5 percentage points around 85 % you need on the order of 400–600 cases. That is why eval suites are large — and why the same cases are reused between versions.

Mastery means

  • Compares two models pairwise on the same questions
  • Decides whether the difference is larger than chance

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences