Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Evals for language models and agents

Be able to build an eval harness with golden answers, LLM-as-judge and regression tests, and to run it in CI.

Prerequisites

Intuition

For code we have tests. For language models and agents we need evals — and they are harder, because the answer is rarely exactly one thing.

Three types, in increasing cost:

  1. Golden answers — exact or normalised match, regex, numeric tolerance. Cheap, deterministic; suits classification, extraction, arithmetic.
  2. Rule-based checks — does it contain a citation? under 200 words? does it avoid giving away the solution? Code that inspects the answer.
  3. LLM-as-judge — a model grades the answer against a rubric, ideally pairwise (A vs B, order randomised). It scales, but the judge has faults of its own: it prefers long answers, its own style, and a position. Calibrate against 100 human judgements; report the judge's agreement.

For agents: measure at the task level (did it solve it?) in a sandbox with a verifiable end state, plus the cost and the number of steps.

In CI: a fixed case set, temperature 0, a threshold plus a regression list, a cost cap per run. AI-grafen's evals/tutor_eval.py is an example: 12 cases checking that the tutor does not give away answers, does not confirm guesses and does not invent facts.

Code

import json, random

RUBRIC = """Grade the ANSWER to the QUESTION using the CONTEXT. Score 1–5 for: correct (supported by the context), complete, concise.
Answer with JSON ONLY: {"correct": n, "complete": n, "concise": n, "reason": "..."}"""

def judge(llm, question, context, answer):
    out = llm(f"{RUBRIC}\n\nQUESTION: {question}\nCONTEXT: {context}\nANSWER: {answer}", temperature=0, json_mode=True)
    return json.loads(out)

def pairwise(llm, question, a, b, rng=random.Random(0)):
    flip = rng.random() < 0.5
    x, y = (b, a) if flip else (a, b)
    out = llm(f"Which answer to '{question}' is better? Answer A or B.\n\nA: {x}\n\nB: {y}", temperature=0).strip()
    winner = "B" if (out.startswith("A")) == flip else "A"       # undo the randomisation of the position
    return winner

def calibrate(judge_scores, human_scores):
    import numpy as np
    return float(np.corrcoef(judge_scores, human_scores)[0, 1])  # report it; < 0.6 → the rubric needs work

# rule-based: the tutor must not give away the solution
def gives_away_solution(answer, key):
    return str(key).lower() in answer.lower()

Mastery means

  • Builds evals with golden answers, LLM-as-judge and regression tests
  • Calibrates the judge against human judgements
  • Runs the evals in CI with a threshold and a cost cap

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences