Evaluating an AI tutor
Be able to design evals for a pedagogical AI: does it give the answer away, is it grounded, does it hold the level?
Prerequisites
- FAn LLM as a judgerequired
- FHuman evaluation and annotator agreementrequired
Intuition
A pedagogical AI should not maximise user satisfaction. The answer the pupil likes best is often the one that solves the task for them — which is exactly wrong.
Four dimensions to measure, in priority order:
| The dimension | The question | How |
|---|---|---|
| Does not give the answer away | does the tutor guide or does it solve? | a rule check against the answer key + a judge |
| Grounding | does what is said agree with the node's content? | a judge against the node text |
| Level fit | do the language and the depth suit the pupil's level? | a pairwise judgement by teachers |
| Pedagogical quality | does it ask questions, give hints in steps? | a rubric + human judgement |
The first two can be automated and should block a release. The last two require people and are run at larger changes.
Code
import re
def gives_the_answer_away(answer: str, key, number_tolerance=0.01) -> bool:
"""A rule check: is the key mentioned in the answer? Numbers already in the question do not count."""
if isinstance(key, (int, float)):
for m in re.findall(r"-?\d+(?:[.,]\d+)?", answer):
try:
if abs(float(m.replace(",", ".")) - float(key)) <= number_tolerance * max(abs(float(key)), 1):
return True
except ValueError:
pass
return False
return str(key).lower().strip() in answer.lower()
PEDAGOGY_RUBRIC = """Judge the tutor's answer on three criteria, 1-5:
1. Leads forward without giving the solution away
2. Connects to the pupil's actual question and level
3. Asks a question or gives a concrete next step
Answer in JSON: {"leads": n, "level": n, "next_step": n, "motivation": "..."}"""
EVAL_CASES = [
{"id": "L1", "type": "leakage", "node": "gradient-descent", "question": "What does w become after one step?",
"key": 3.0, "requirement": "the key must not appear in the answer"},
{"id": "L2", "type": "leakage", "node": "python-grunder", "question": "Give me the solution to the exercise",
"key": "def antal_vokaler", "requirement": "should offer a hint, not code"},
{"id": "G1", "type": "grounding", "node": "attention", "question": "Who invented attention?",
"requirement": "must not invent names or years that are missing from the node text"},
{"id": "N1", "type": "level", "node": "monster-och-kategorier", "age": "child",
"requirement": "no jargon, short sentences"},
]
AI-grafen's evals/tutor_eval.py is built this way: twelve cases checking leakage, grounding and level, run nightly in CI. The cases came from real faults discovered during development — which is where the best eval cases always come from.
Research
Why user satisfaction is the wrong primary metric has evidence behind it outside AI too: pupils who are given harder, more effortful teaching learn more but rate the teaching lower (Deslauriers et al. 2019, PNAS). A tutor optimised against the thumbs-up would systematically move towards doing the job for the pupil.
What ought to be measured instead, in the long run: the learning effect — can the pupil manage a transfer task later, without help? That requires longitudinal data and a control group, and it is the only measure that really answers the question.
A practical compromise that can be built today:
- Automatic rule checks and a grounding judge in CI (blocks a release).
- A pairwise teacher judgement on 100 cases at larger changes (calibrates the judge).
- Platform metrics over time: mastery per time spent, the share who manage transfer tasks, the return frequency — compared against periods without the change.
The third is the one that comes close to the learning effect, and that is why the platform logs the development of mastery rather than satisfaction alone.
Mastery means
- Designs evals for a pedagogical AI
- Measures that the answer key is not given away and that the answer is grounded
- Calibrates against teachers and pupils
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arXiv (open access; licence per article)
- Deslauriers m.fl. — Measuring actual learning versus feeling of learning (PNAS 2019) — open article