Skip to content
AI-grafen
EUniversityEvals and benchmarks· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Model evaluation

Be able to design an evaluation with the right metrics, keep the training and test data apart, and interpret confidence intervals.

Prerequisites

Intuition

An evaluation answers one question: is model A better than B for task X, measured as Y, on data that resembles production? Every part of that sentence is a design choice.

  • Task/data: the test set should resemble what the model meets — not what was easy to collect.
  • Metric: tied to usefulness (F1 for imbalance, MAE in kronor, «solved the task» for agents). One metric is rarely enough; report 2–3 and say which one decides.
  • Uncertainty: n = 200 gives ±5 percentage points. Without an interval the number is an anecdote.
  • Comparison: the same test cases for every model → a paired comparison (wins/losses per case) is far more sensitive than two separate means.

Common mistakes: test data leaked into the training; picking the best of ten runs; switching metric when the first one looked bad; evaluating on data from last year.

Code

import numpy as np

def bootstrap_ci(x, stat=np.mean, n=2000, alpha=0.05, seed=0):
    rng = np.random.default_rng(seed); x = np.asarray(x)
    s = [stat(rng.choice(x, len(x), replace=True)) for _ in range(n)]
    return stat(x), np.percentile(s, 100 * alpha / 2), np.percentile(s, 100 * (1 - alpha / 2))

correct_a = np.array([1,1,0,1,1,1,0,1,1,1,0,1,1,1,1,0,1,1,1,1])   # per test case
correct_b = np.array([1,0,0,1,1,1,0,1,0,1,0,1,1,1,1,0,1,0,1,1])
print(bootstrap_ci(correct_a))               # (0.8, 0.6, 0.95)
print(bootstrap_ci(correct_b))               # (0.7, 0.5, 0.9)  — the intervals overlap

# the paired comparison: the same cases
diff = correct_a - correct_b
print(bootstrap_ci(diff))                    # (0.1, -0.05, 0.25) → not settled with n=20
print("A won", int((diff == 1).sum()), "B won", int((diff == -1).sum()))

For LLM output with no answer key: LLM-as-judge with a clear rubric, pairwise (A vs B, order randomised), and a check of the judge against human judgements on a sample.

Mastery means

  • Designs an evaluation with the right metrics and separate test data
  • Reports confidence intervals and compares models pairwise
  • Recognises the common mistakes: leakage, cherry-picking, the wrong metric

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences