Training data for language models: filtering and dedup
Be able to build a filtered, deduplicated text corpus and measure its quality.
Prerequisites
- EWeb scraping — the technique and the rulesrequired
- FDataset design for fine-tuningrequired
Intuition
Raw web text is for the most part unusable: boilerplate, navigation menus, auto-generated text, spam and duplicates. The filtering is what decides the model's quality — more than the architecture does.
The pipeline, in order:
| Step | Typically discards |
|---|---|
| 1. Extract the text from the HTML | tags, scripts, the menu |
| 2. Language identification | the wrong language |
| 3. Quality heuristics | too short, too many symbols, word lists |
| 4. Deduplication | 30–60 % of everything |
| 5. Toxicity and PII filters | harmful content, personal data |
| 6. A contamination check | test data that has leaked in |
Step 4 is the most impactful. Lee et al. (2021) showed that deduplication both reduces memorisation and improves the model — fewer repetitions means the model does not learn the same text a hundred times.
A rule of thumb that holds: ten times less but cleaned data beats ten times more uncleaned.
Formal
Deduplication at three levels:
| Level | Finds | Method |
|---|---|---|
| Exact | identical documents | a hash of the normalised text |
| Near-duplicates | the same text with small changes | MinHash + LSH |
| Substring level | recurring passages in different documents | a suffix array |
MinHash approximates the Jaccard similarity between shingle sets, and LSH makes it possible to find the candidates without comparing every pair — without it, dedup at billion scale is impossible.
Quality heuristics (from the Gopher and C4 work) — simple rules that remove a lot of rubbish:
| Rule | Discards |
|---|---|
| < 50 or > 100 000 words | fragments and dumps |
| A mean word length outside 3–10 characters | code, noise, tables |
| > 90 % of the lines start with a bullet | navigation |
| < 80 % of the words in a word list | noise and code litter |
| A symbol share > 10 % | scripts, tables |
| No punctuation | lists and menus |
Model-based filtering goes one step further: train a classifier on «good» text (Wikipedia, books) against random web text and keep what resembles the former. Effective, but it homogenises the corpus and risks systematically filtering out dialects, minority languages and informal registers.
Contamination is the check most often missing. If the test set has happened to end up in the training data, every evaluation is meaningless. Check it with n-gram overlap between the corpus and every benchmark you intend to use, and report the result.
Personal data in the corpus requires an explicit position: detect and mask emails, phone numbers, national ID numbers and addresses. That is not enough to make the corpus GDPR-safe, but it removes the most obvious.
Measure the corpus, not just its size:
| Metric | Tells you |
|---|---|
| Total and unique tokens | the volume and the variation |
| The share remaining after each filter step | where the data disappears |
| The perplexity under a reference model | rubbish in the tails |
| The domain distribution | what the model will know |
| The language distribution | particularly important for Swedish |
| The contamination rate per benchmark | whether the evaluation holds |
Code
import hashlib, re, unicodedata
from collections import Counter
def normalise(text: str) -> str:
t = unicodedata.normalize("NFKC", text).lower()
return re.sub(r"\s+", " ", t).strip()
def exact_dedup(documents):
seen, out = set(), []
for d in documents:
h = hashlib.sha256(normalise(d).encode()).hexdigest()
if h not in seen:
seen.add(h); out.append(d)
return out
# MinHash + LSH for near-duplicates
def shingles(text, n=5):
words = normalise(text).split()
return {" ".join(words[i:i + n]) for i in range(max(len(words) - n + 1, 1))}
def minhash(s, count=128, seed=0):
import random
rng = random.Random(seed)
seeds = [rng.getrandbits(64) for _ in range(count)]
sig = []
for f in seeds:
sig.append(min(hash((sh, f)) & 0xFFFFFFFF for sh in s) if s else 0)
return tuple(sig)
def lsh_bands(sig, bands=16):
r = len(sig) // bands
return [hash(sig[i * r:(i + 1) * r]) for i in range(bands)]
def near_dedup(documents, bands=16, threshold=0.8):
buckets, keep = {}, []
for i, d in enumerate(documents):
s = shingles(d)
sig = minhash(s)
candidates = set()
keys = lsh_bands(sig, bands)
for key in keys:
candidates |= buckets.get(key, set())
duplicate = any(
len(s & shingles(documents[j])) / max(len(s | shingles(documents[j])), 1) > threshold
for j in candidates)
if not duplicate:
keep.append(d)
for key in keys:
buckets.setdefault(key, set()).add(i)
return keep
# Quality heuristics
SYMBOLS = set("#<>{}[]|\\@^~`")
def quality_ok(text, wordlist=None):
words = text.split()
if not 50 <= len(words) <= 100_000:
return False, "length"
if not 3 <= sum(len(w) for w in words) / len(words) <= 10:
return False, "mean word length"
lines = [l for l in text.splitlines() if l.strip()]
if lines and sum(l.lstrip().startswith(("•", "-", "*")) for l in lines) / len(lines) > 0.9:
return False, "bullet list"
if sum(c in SYMBOLS for c in text) / max(len(text), 1) > 0.10:
return False, "symbols"
if not re.search(r"[.!?]", text):
return False, "no punctuation"
if wordlist and sum(w.strip(".,!?").lower() in wordlist for w in words) / len(words) < 0.80:
return False, "word list"
return True, "ok"
# A contamination check against the benchmarks
def contamination(corpus_ngrams: set, benchmark_texts, n=13):
hits = 0
for t in benchmark_texts:
words = normalise(t).split()
grams = {" ".join(words[i:i + n]) for i in range(max(len(words) - n + 1, 1))}
if grams & corpus_ngrams:
hits += 1
return {"contaminated": hits, "of": len(benchmark_texts),
"share": round(hits / max(len(benchmark_texts), 1), 4)}
# Mask the personal data
PATTERNS = {
"EMAIL": r"\b[\w.+-]+@[\w-]+\.[\w.]+\b",
"PHONE": r"\b0[\d\s-]{7,12}\b",
"NATIONAL_ID": r"\b(19|20)?\d{6}[-+]?\d{4}\b",
}
def mask_pii(text):
counts = Counter()
for name, m in PATTERNS.items():
text, n = re.subn(m, f"<{name}>", text)
counts[name] += n
return text, dict(counts)
# Report where the data disappears
def pipeline_report(documents, wordlist=None):
steps = {"in": len(documents)}
d = exact_dedup(documents); steps["after exact dedup"] = len(d)
d = [x for x in d if quality_ok(x, wordlist)[0]]; steps["after quality"] = len(d)
d = near_dedup(d); steps["after near-dedup"] = len(d)
steps["share_remaining"] = round(len(d) / max(len(documents), 1), 3)
return d, steps
Mastery means
- Builds a filtering pipeline
- Deduplicates at several levels
- Measures corpus quality and contamination
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Deduplicating Training Data Makes Language Models Better — arXiv (open access; licence per article)
- arXiv — Scaling Language Models: Methods, Analysis & Insights from Training Gopher — arXiv (open access; licence per article)
- arXiv — The RefinedWeb Dataset for Falcon LLM — arXiv (open access; licence per article)