EUniversityEvals and benchmarks· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Model evaluation
Be able to design an evaluation with the right metrics, keep the training and test data apart, and interpret confidence intervals.
Prerequisites
Intuition
An evaluation answers one question: is model A better than B for task X, measured as Y, on data that resembles production? Every part of that sentence is a design choice.
- Task/data: the test set should resemble what the model meets — not what was easy to collect.
- Metric: tied to usefulness (F1 for imbalance, MAE in kronor, «solved the task» for agents). One metric is rarely enough; report 2–3 and say which one decides.
- Uncertainty: n = 200 gives ±5 percentage points. Without an interval the number is an anecdote.
- Comparison: the same test cases for every model → a paired comparison (wins/losses per case) is far more sensitive than two separate means.
Common mistakes: test data leaked into the training; picking the best of ten runs; switching metric when the first one looked bad; evaluating on data from last year.
Code
import numpy as np
def bootstrap_ci(x, stat=np.mean, n=2000, alpha=0.05, seed=0):
rng = np.random.default_rng(seed); x = np.asarray(x)
s = [stat(rng.choice(x, len(x), replace=True)) for _ in range(n)]
return stat(x), np.percentile(s, 100 * alpha / 2), np.percentile(s, 100 * (1 - alpha / 2))
correct_a = np.array([1,1,0,1,1,1,0,1,1,1,0,1,1,1,1,0,1,1,1,1]) # per test case
correct_b = np.array([1,0,0,1,1,1,0,1,0,1,0,1,1,1,1,0,1,0,1,1])
print(bootstrap_ci(correct_a)) # (0.8, 0.6, 0.95)
print(bootstrap_ci(correct_b)) # (0.7, 0.5, 0.9) — the intervals overlap
# the paired comparison: the same cases
diff = correct_a - correct_b
print(bootstrap_ci(diff)) # (0.1, -0.05, 0.25) → not settled with n=20
print("A won", int((diff == 1).sum()), "B won", int((diff == -1).sum()))
For LLM output with no answer key: LLM-as-judge with a clear rubric, pairwise (A vs B, order randomised), and a check of the judge against human judgements on a sample.
Mastery means
- Designs an evaluation with the right metrics and separate test data
- Reports confidence intervals and compares models pairwise
- Recognises the common mistakes: leakage, cherry-picking, the wrong metric
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Bootstrapping (statistics) (CC BY-SA 4.0) — CC BY-SA 4.0
- arXiv — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arXiv (open access; licence per article)
Leads to
Part of the goals (18)
- Fine-tune a model with LoRA
- Build a RAG system you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- Frontier Lab — an independent research project
- Fine-tune and run your own models
- Build a memory system for an agent
- Build an agent you can trust
- AI safety in practice
- Evals in practice
- Responsible AI in practice
- Build an AI service that survives production
- An AI service in operation
- Reproduce a paper
- Deep reinforcement learning
- Interpreting a language model
- Statistics for experiments