EUniversityEvals and benchmarks· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Build an eval harness
Be able to build a harness with test cases, metrics and a report that runs in CI.
Prerequisites
- DTesting with pytestrequired
- EModel evaluationrequired
Intuition
An eval harness is to models what unit tests are to code: a fixed set of cases with expected answers, a script that runs the model against them, and a report. Every change (the prompt, the model version, a retrieval parameter) is run through the harness before it ships.
The parts:
- Cases:
{id, input, expected, kind}in JSONL, under version control. Mix easy ones, hard ones, edge cases and previous bugs. - A metric per case: exact match, normalised match, numeric tolerance, a rubric via a judge.
- A report: the overall metric plus a list of the cases that went from right to wrong (regressions) — the most important part.
- Determinism: temperature 0, fixed seeds, sorted output; otherwise the report cannot be compared.
- A CI gate: the build fails if the metric falls below the threshold or if a «gold standard» case has regressed.
Code
import json
def load_cases(path):
return [json.loads(l) for l in open(path, encoding="utf-8") if l.strip()]
def score(pred, exp, kind):
if kind == "exact": return pred == exp
if kind == "normalized": return " ".join(pred.lower().split()) == " ".join(exp.lower().split())
if kind == "numeric": return abs(float(pred) - float(exp)) <= 1e-6
raise ValueError(kind)
def run_eval(model_fn, cases):
res = {c["id"]: score(model_fn(c["input"]), c["expected"], c["kind"]) for c in cases}
fails = sorted(k for k, ok in res.items() if not ok)
return {"n": len(cases), "accuracy": sum(res.values()) / len(cases), "failures": fails, "per_case": res}
def compare(before, after):
return sorted(k for k in before["per_case"] if before["per_case"][k] and not after["per_case"].get(k, False))
# In CI:
# r = run_eval(model, load_cases("cases.jsonl")); assert r["accuracy"] >= 0.85, r["failures"]
# reg = compare(json.load(open("baseline.json")), r); assert not reg, f"regressions: {reg}"
The lab evals-harness-labb builds exactly this, with tests.
Mastery means
- Builds a harness with test cases, metrics and a deterministic report
- Finds regressions between two versions
- Runs the harness in CI with a threshold
Sign in to do the exercises and build your mastery up.