Skip to content
AI-grafen
EUniversityEvals and benchmarks· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Benchmarks: what they measure and miss

Be able to read MMLU, HumanEval and similar results critically.

Prerequisites

Intuition

The benchmarkMeasuresThe pitfall
MMLUmultiple choice in 57 academic subjectsmemorisation, sensitive to the prompt format and n-shot
GSM8K / MATHmathematical word problemscontamination, chain-of-thought changes everything
HumanEval / MBPPPython functions against testssmall, well-known problems — saturated
MT-Benchmulti-turn dialogues judged by GPT-4the judge's bias (length, style)
Chatbot Arenahuman pairwise preferences (Elo)it measures «is liked», not «is right»
SWE-benchsolving real GitHub issuesexpensive, the agent scaffold matters more than the model

Always read: which version, n-shot, the prompt, how the answer was parsed, whether the test data was on the web before the model's training data (contamination). Two percentage points on MMLU says nothing about your customer service. The only benchmark that certainly applies to your task is the one you build yourselves.

Code

# Report reproducibly: the exact configuration
config = {
  "benchmark": "gsm8k", "split": "test", "n": 1319,
  "shots": 8, "prompt_style": "chain-of-thought", "answer_parse": "regex r'#### (\\-?\\d+)'",
  "model": "qwen2.5-7b-instruct", "temperature": 0, "max_tokens": 512,
  "seed": 0, "harness": "lm-eval 0.4.5",
}
# With a different parse regex the same model can come out 5 pp differently.

# A contamination check (rough): n-gram overlap between the test questions and the training corpus
def ngram_overlap(test, corpus_ngrams, n=13):
    toks = test.split()
    grams = {" ".join(toks[i:i+n]) for i in range(len(toks) - n + 1)}
    return len(grams & corpus_ngrams) / max(1, len(grams))

Mastery means

  • Describes what MMLU, HumanEval, GSM8K, MT-Bench and the Arena measure
  • Reads benchmark results critically: contamination, the prompt format, n-shot
  • Decides which benchmark is relevant for a given task

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences