FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Evals for language models and agents
Be able to build an eval harness with golden answers, LLM-as-judge and regression tests, and to run it in CI.
Prerequisites
- EModel evaluationrequired
- ERAG — retrieval-augmented generationhelpful
Intuition
For code we have tests. For language models and agents we need evals — and they are harder, because the answer is rarely exactly one thing.
Three types, in increasing cost:
- Golden answers — exact or normalised match, regex, numeric tolerance. Cheap, deterministic; suits classification, extraction, arithmetic.
- Rule-based checks — does it contain a citation? under 200 words? does it avoid giving away the solution? Code that inspects the answer.
- LLM-as-judge — a model grades the answer against a rubric, ideally pairwise (A vs B, order randomised). It scales, but the judge has faults of its own: it prefers long answers, its own style, and a position. Calibrate against 100 human judgements; report the judge's agreement.
For agents: measure at the task level (did it solve it?) in a sandbox with a verifiable end state, plus the cost and the number of steps.
In CI: a fixed case set, temperature 0, a threshold plus a regression list, a cost cap per run. AI-grafen's evals/tutor_eval.py is an example: 12 cases checking that the tutor does not give away answers, does not confirm guesses and does not invent facts.
Code
import json, random
RUBRIC = """Grade the ANSWER to the QUESTION using the CONTEXT. Score 1–5 for: correct (supported by the context), complete, concise.
Answer with JSON ONLY: {"correct": n, "complete": n, "concise": n, "reason": "..."}"""
def judge(llm, question, context, answer):
out = llm(f"{RUBRIC}\n\nQUESTION: {question}\nCONTEXT: {context}\nANSWER: {answer}", temperature=0, json_mode=True)
return json.loads(out)
def pairwise(llm, question, a, b, rng=random.Random(0)):
flip = rng.random() < 0.5
x, y = (b, a) if flip else (a, b)
out = llm(f"Which answer to '{question}' is better? Answer A or B.\n\nA: {x}\n\nB: {y}", temperature=0).strip()
winner = "B" if (out.startswith("A")) == flip else "A" # undo the randomisation of the position
return winner
def calibrate(judge_scores, human_scores):
import numpy as np
return float(np.corrcoef(judge_scores, human_scores)[0, 1]) # report it; < 0.6 → the rubric needs work
# rule-based: the tutor must not give away the solution
def gives_away_solution(answer, key):
return str(key).lower() in answer.lower()
Mastery means
- Builds evals with golden answers, LLM-as-judge and regression tests
- Calibrates the judge against human judgements
- Runs the evals in CI with a threshold and a cost cap
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arXiv (open access; licence per article)
- OpenAI Evals (MIT) — MIT