Benchmark contamination
Be able to detect and avoid test data leaking into the training.
Prerequisites
- DTraining, validation and testrequired
- EBenchmarks: what they measure and missrequired
Intuition
If a model's training data contains the test questions (with the answers) the benchmark measures memory, not ability. Web-scraped corpora contain GitHub, StackOverflow, Wikipedia — and GSM8K, HumanEval and MMLU lie there in plain text.
The routes in: public benchmarks on the web; synthetic training data generated by a model that has seen the benchmark; fine-tuning data that happens to contain test cases; «training on the validation set» by mistake.
Detection:
- n-gram overlap between the test cases and the training corpus (13-grams are the standard) — it requires access to the corpus.
- Perplexity: abnormally low on the test set compared with newly written text in the same style.
- A canary test: the model completes a test question verbatim when it is given the beginning.
- A time comparison: the performance on cases published after the model's training date.
The protection: private evals that are never put on the web; newly written cases; canary strings in your own datasets; document the date.
Code
import re
def ngrams(text, n=13):
t = re.findall(r"\w+", text.lower())
return {" ".join(t[i:i+n]) for i in range(len(t) - n + 1)}
def contaminated(test_case, corpus_ngrams, n=13, thr=0.0):
g = ngrams(test_case, n)
return len(g & corpus_ngrams) / max(1, len(g)) > thr
# The canary test: give the first half of the question, see whether the model continues verbatim
def canary(llm, question):
words = question.split(); half = " ".join(words[: len(words) // 2])
continuation = llm(half, max_tokens=len(words), temperature=0)
return " ".join(words[len(words) // 2:]).lower() in continuation.lower()
# The perplexity comparison: PPL(test) << PPL(newly written cases in the same style) → suspicious
Always report: «the model's training data cut-off» and «the benchmark's publication date». If the cut-off > the publication: assume contamination until the opposite is shown.
Mastery means
- Explains how test data leaks into pre-training and fine-tuning
- Detects contamination with n-gram overlap and a perplexity comparison
- Designs uncontaminated evals
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Investigating Data Contamination in Modern Benchmarks for LLMs — arXiv (open access; licence per article)
- arXiv — Language Models are Few-Shot Learners (appendix C, kontaminering) — arXiv (open access; licence per article)