DAI developerScientific method· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN
Measure the AI fairly
Be able to design a fair test of two AI tools with the same questions and the same answer key.
Prerequisites
- BA fair experimentrequired
- CEvaluate AI performance with your own questionsrequired
Intuition
«Tool A is better than B» is only a claim until it has been measured fairly:
- The same questions to both — at least 30, preferably 100, mixed difficulty, written before you test.
- An answer key written in advance, by you, not by either of the tools.
- The same conditions: the same wording, no follow-up questions, the same day.
- Blind marking: you should not know which tool gave the answer while you mark it.
- Report the share correct and the uncertainty (±) and what kinds of mistake each tool made.
30 questions give ±8 percentage points — differences smaller than that are noise.
Interactive
A template:
| # | Question | Key | A | B |
|---|---|---|---|---|
| 1 | … | … | ✔ | ✘ |
Result: A 24/30 (80 % ± 7), B 21/30 (70 % ± 8). They overlap → not settled; but look at which questions: B missed 5 of the 6 arithmetic questions, A missed years. That is more useful than the total.
A paired comparison: count the cases where A is right & B wrong (7) against B right & A wrong (4). 7 against 4 out of 11 — still weak. The honest answer is often «we need more questions».
Mastery means
- Designs a fair test of two AI tools with the same questions and answer key
- Reports the result with its uncertainty and with error categories
Sign in to do the exercises and build your mastery up.
Sources
- Skolverket — About AI in school (in Swedish) — Skolverket's open terms
- Wikipedia — A/B testing (CC BY-SA 4.0) — CC BY-SA 4.0