EUniversityLab· about 60 min· server sandbox
Lab: build an eval harness
Build a harness that runs a model (here: a rule-based function) against test cases with an answer key, computes metrics, finds regressions and writes a report.
Teaches: Build an eval harness
Requires: Model evaluationTesting with pytest
Theory
An eval is a list of cases {input, expected} + a metric (exact match, normalised match, numeric tolerance). The harness should be deterministic and comparable between versions.
Sub-tasks
- load_cases —
load_cases(path)reads JSONL into a list of dicts. - score —
score(pred, expected, kind)for kind exact|normalized|numeric (tol 1e-6). - run_eval —
run_eval(model_fn, cases)→ {accuracy, n, failures:[…]};compare(a, b)→ regressions (cases that went from right to wrong).
Passes when: accuracy >= 0.8
The starter code
runs in an isolated sandbox on the serverimport json
def baseline(q):
"""Regelbaserad 'modell' att utvärdera."""
q = q.lower()
if "huvudstad" in q and "sverige" in q:
return "Stockholm"
if q.startswith("vad är") and "+" in q:
a, b = q.replace("vad är", "").replace("?", "").split("+")
return str(int(a) + int(b))
return "vet inte"
def load_cases(path):
# TODO: en dict per rad
...
def score(pred, expected, kind="exact"):
# TODO: exact | normalized (lower + strip + ett mellanslag) | numeric (abs diff < 1e-6)
...
def run_eval(model_fn, cases):
# TODO: {"accuracy": andel rätt, "n": antal, "failures": [{"input", "expected", "got"}], "results": {input: bool}}
...
def compare(before, after):
# TODO: lista inputs där before rätt och after fel
...
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
The baseline model (baseline) scores ≥ 0.8 on the cases; compare finds exactly the cases that regressed.
Common mistakes
- Does not normalise whitespace/case in 'normalized'.
- Compares floats with ==.
- The report is not deterministic (set ordering).