Skip to content
AI-grafen
DAI developerScientific method· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Measure the AI fairly

Be able to design a fair test of two AI tools with the same questions and the same answer key.

Prerequisites

Intuition

«Tool A is better than B» is only a claim until it has been measured fairly:

  1. The same questions to both — at least 30, preferably 100, mixed difficulty, written before you test.
  2. An answer key written in advance, by you, not by either of the tools.
  3. The same conditions: the same wording, no follow-up questions, the same day.
  4. Blind marking: you should not know which tool gave the answer while you mark it.
  5. Report the share correct and the uncertainty (±) and what kinds of mistake each tool made.

30 questions give ±8 percentage points — differences smaller than that are noise.

Interactive

A template:

#QuestionKeyAB
1……✔✘

Result: A 24/30 (80 % ± 7), B 21/30 (70 % ± 8). They overlap → not settled; but look at which questions: B missed 5 of the 6 arithmetic questions, A missed years. That is more useful than the total.

A paired comparison: count the cases where A is right & B wrong (7) against B right & A wrong (4). 7 against 4 out of 11 — still weak. The honest answer is often «we need more questions».

Mastery means

  • Designs a fair test of two AI tools with the same questions and answer key
  • Reports the result with its uncertainty and with error categories

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences