Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Benchmark contamination

Be able to detect and avoid test data leaking into the training.

Prerequisites

Intuition

If a model's training data contains the test questions (with the answers) the benchmark measures memory, not ability. Web-scraped corpora contain GitHub, StackOverflow, Wikipedia — and GSM8K, HumanEval and MMLU lie there in plain text.

The routes in: public benchmarks on the web; synthetic training data generated by a model that has seen the benchmark; fine-tuning data that happens to contain test cases; «training on the validation set» by mistake.

Detection:

  • n-gram overlap between the test cases and the training corpus (13-grams are the standard) — it requires access to the corpus.
  • Perplexity: abnormally low on the test set compared with newly written text in the same style.
  • A canary test: the model completes a test question verbatim when it is given the beginning.
  • A time comparison: the performance on cases published after the model's training date.

The protection: private evals that are never put on the web; newly written cases; canary strings in your own datasets; document the date.

Code

import re

def ngrams(text, n=13):
    t = re.findall(r"\w+", text.lower())
    return {" ".join(t[i:i+n]) for i in range(len(t) - n + 1)}

def contaminated(test_case, corpus_ngrams, n=13, thr=0.0):
    g = ngrams(test_case, n)
    return len(g & corpus_ngrams) / max(1, len(g)) > thr

# The canary test: give the first half of the question, see whether the model continues verbatim
def canary(llm, question):
    words = question.split(); half = " ".join(words[: len(words) // 2])
    continuation = llm(half, max_tokens=len(words), temperature=0)
    return " ".join(words[len(words) // 2:]).lower() in continuation.lower()

# The perplexity comparison: PPL(test) << PPL(newly written cases in the same style) → suspicious

Always report: «the model's training data cut-off» and «the benchmark's publication date». If the cut-off > the publication: assume contamination until the opposite is shown.

Mastery means

  • Explains how test data leaks into pre-training and fine-tuning
  • Detects contamination with n-gram overlap and a perplexity comparison
  • Designs uncontaminated evals

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences