Skip to content
AI-grafen
EUniversityLab· about 60 min· server sandbox

Lab: build an eval harness

Build a harness that runs a model (here: a rule-based function) against test cases with an answer key, computes metrics, finds regressions and writes a report.

Theory

An eval is a list of cases {input, expected} + a metric (exact match, normalised match, numeric tolerance). The harness should be deterministic and comparable between versions.

Sub-tasks

  1. load_cases — load_cases(path) reads JSONL into a list of dicts.
  2. score — score(pred, expected, kind) for kind exact|normalized|numeric (tol 1e-6).
  3. run_eval — run_eval(model_fn, cases) → {accuracy, n, failures:[…]}; compare(a, b) → regressions (cases that went from right to wrong).

Passes when: accuracy >= 0.8

The starter code

runs in an isolated sandbox on the server
import json


def baseline(q):
    """Regelbaserad 'modell' att utvärdera."""
    q = q.lower()
    if "huvudstad" in q and "sverige" in q:
        return "Stockholm"
    if q.startswith("vad är") and "+" in q:
        a, b = q.replace("vad är", "").replace("?", "").split("+")
        return str(int(a) + int(b))
    return "vet inte"


def load_cases(path):
    # TODO: en dict per rad
    ...


def score(pred, expected, kind="exact"):
    # TODO: exact | normalized (lower + strip + ett mellanslag) | numeric (abs diff < 1e-6)
    ...


def run_eval(model_fn, cases):
    # TODO: {"accuracy": andel rätt, "n": antal, "failures": [{"input", "expected", "got"}], "results": {input: bool}}
    ...


def compare(before, after):
    # TODO: lista inputs där before rätt och after fel
    ...

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

The baseline model (baseline) scores ≥ 0.8 on the cases; compare finds exactly the cases that regressed.

Common mistakes

  • Does not normalise whitespace/case in 'normalized'.
  • Compares floats with ==.
  • The report is not deterministic (set ordering).