Skip to content
AI-grafen
EUniversityEvals and benchmarks· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Build an eval harness

Be able to build a harness with test cases, metrics and a report that runs in CI.

Prerequisites

Intuition

An eval harness is to models what unit tests are to code: a fixed set of cases with expected answers, a script that runs the model against them, and a report. Every change (the prompt, the model version, a retrieval parameter) is run through the harness before it ships.

The parts:

  • Cases: {id, input, expected, kind} in JSONL, under version control. Mix easy ones, hard ones, edge cases and previous bugs.
  • A metric per case: exact match, normalised match, numeric tolerance, a rubric via a judge.
  • A report: the overall metric plus a list of the cases that went from right to wrong (regressions) — the most important part.
  • Determinism: temperature 0, fixed seeds, sorted output; otherwise the report cannot be compared.
  • A CI gate: the build fails if the metric falls below the threshold or if a «gold standard» case has regressed.

Code

import json

def load_cases(path):
    return [json.loads(l) for l in open(path, encoding="utf-8") if l.strip()]

def score(pred, exp, kind):
    if kind == "exact": return pred == exp
    if kind == "normalized": return " ".join(pred.lower().split()) == " ".join(exp.lower().split())
    if kind == "numeric": return abs(float(pred) - float(exp)) <= 1e-6
    raise ValueError(kind)

def run_eval(model_fn, cases):
    res = {c["id"]: score(model_fn(c["input"]), c["expected"], c["kind"]) for c in cases}
    fails = sorted(k for k, ok in res.items() if not ok)
    return {"n": len(cases), "accuracy": sum(res.values()) / len(cases), "failures": fails, "per_case": res}

def compare(before, after):
    return sorted(k for k in before["per_case"] if before["per_case"][k] and not after["per_case"].get(k, False))

# In CI:
# r = run_eval(model, load_cases("cases.jsonl")); assert r["accuracy"] >= 0.85, r["failures"]
# reg = compare(json.load(open("baseline.json")), r); assert not reg, f"regressions: {reg}"

The lab evals-harness-labb builds exactly this, with tests.

Mastery means

  • Builds a harness with test cases, metrics and a deterministic report
  • Finds regressions between two versions
  • Runs the harness in CI with a threshold

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences