Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Evaluating an AI tutor

Be able to design evals for a pedagogical AI: does it give the answer away, is it grounded, does it hold the level?

Prerequisites

Intuition

A pedagogical AI should not maximise user satisfaction. The answer the pupil likes best is often the one that solves the task for them — which is exactly wrong.

Four dimensions to measure, in priority order:

The dimensionThe questionHow
Does not give the answer awaydoes the tutor guide or does it solve?a rule check against the answer key + a judge
Groundingdoes what is said agree with the node's content?a judge against the node text
Level fitdo the language and the depth suit the pupil's level?a pairwise judgement by teachers
Pedagogical qualitydoes it ask questions, give hints in steps?a rubric + human judgement

The first two can be automated and should block a release. The last two require people and are run at larger changes.

Code

import re

def gives_the_answer_away(answer: str, key, number_tolerance=0.01) -> bool:
    """A rule check: is the key mentioned in the answer? Numbers already in the question do not count."""
    if isinstance(key, (int, float)):
        for m in re.findall(r"-?\d+(?:[.,]\d+)?", answer):
            try:
                if abs(float(m.replace(",", ".")) - float(key)) <= number_tolerance * max(abs(float(key)), 1):
                    return True
            except ValueError:
                pass
        return False
    return str(key).lower().strip() in answer.lower()

PEDAGOGY_RUBRIC = """Judge the tutor's answer on three criteria, 1-5:
1. Leads forward without giving the solution away
2. Connects to the pupil's actual question and level
3. Asks a question or gives a concrete next step
Answer in JSON: {"leads": n, "level": n, "next_step": n, "motivation": "..."}"""

EVAL_CASES = [
    {"id": "L1", "type": "leakage", "node": "gradient-descent", "question": "What does w become after one step?",
     "key": 3.0, "requirement": "the key must not appear in the answer"},
    {"id": "L2", "type": "leakage", "node": "python-grunder", "question": "Give me the solution to the exercise",
     "key": "def antal_vokaler", "requirement": "should offer a hint, not code"},
    {"id": "G1", "type": "grounding", "node": "attention", "question": "Who invented attention?",
     "requirement": "must not invent names or years that are missing from the node text"},
    {"id": "N1", "type": "level", "node": "monster-och-kategorier", "age": "child",
     "requirement": "no jargon, short sentences"},
]

AI-grafen's evals/tutor_eval.py is built this way: twelve cases checking leakage, grounding and level, run nightly in CI. The cases came from real faults discovered during development — which is where the best eval cases always come from.

Research

Why user satisfaction is the wrong primary metric has evidence behind it outside AI too: pupils who are given harder, more effortful teaching learn more but rate the teaching lower (Deslauriers et al. 2019, PNAS). A tutor optimised against the thumbs-up would systematically move towards doing the job for the pupil.

What ought to be measured instead, in the long run: the learning effect — can the pupil manage a transfer task later, without help? That requires longitudinal data and a control group, and it is the only measure that really answers the question.

A practical compromise that can be built today:

  1. Automatic rule checks and a grounding judge in CI (blocks a release).
  2. A pairwise teacher judgement on 100 cases at larger changes (calibrates the judge).
  3. Platform metrics over time: mastery per time spent, the share who manage transfer tasks, the return frequency — compared against periods without the change.

The third is the one that comes close to the learning effect, and that is why the platform logs the development of mastery rather than satisfaction alone.

Mastery means

  • Designs evals for a pedagogical AI
  • Measures that the answer key is not given away and that the answer is grounded
  • Calibrates against teachers and pupils

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences