EUniversityEvals and benchmarks· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Benchmarks: what they measure and miss
Be able to read MMLU, HumanEval and similar results critically.
Prerequisites
- EModel evaluationrequired
Intuition
| The benchmark | Measures | The pitfall |
|---|---|---|
| MMLU | multiple choice in 57 academic subjects | memorisation, sensitive to the prompt format and n-shot |
| GSM8K / MATH | mathematical word problems | contamination, chain-of-thought changes everything |
| HumanEval / MBPP | Python functions against tests | small, well-known problems — saturated |
| MT-Bench | multi-turn dialogues judged by GPT-4 | the judge's bias (length, style) |
| Chatbot Arena | human pairwise preferences (Elo) | it measures «is liked», not «is right» |
| SWE-bench | solving real GitHub issues | expensive, the agent scaffold matters more than the model |
Always read: which version, n-shot, the prompt, how the answer was parsed, whether the test data was on the web before the model's training data (contamination). Two percentage points on MMLU says nothing about your customer service. The only benchmark that certainly applies to your task is the one you build yourselves.
Code
# Report reproducibly: the exact configuration
config = {
"benchmark": "gsm8k", "split": "test", "n": 1319,
"shots": 8, "prompt_style": "chain-of-thought", "answer_parse": "regex r'#### (\\-?\\d+)'",
"model": "qwen2.5-7b-instruct", "temperature": 0, "max_tokens": 512,
"seed": 0, "harness": "lm-eval 0.4.5",
}
# With a different parse regex the same model can come out 5 pp differently.
# A contamination check (rough): n-gram overlap between the test questions and the training corpus
def ngram_overlap(test, corpus_ngrams, n=13):
toks = test.split()
grams = {" ".join(toks[i:i+n]) for i in range(len(toks) - n + 1)}
return len(grams & corpus_ngrams) / max(1, len(grams))
Mastery means
- Describes what MMLU, HumanEval, GSM8K, MT-Bench and the Arena measure
- Reads benchmark results critically: contamination, the prompt format, n-shot
- Decides which benchmark is relevant for a given task
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Measuring Massive Multitask Language Understanding — arXiv (open access; licence per article)
- arXiv — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arXiv (open access; licence per article)
- Hugging Face Open LLM Leaderboard — free to read