Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 90 min· fast-moving, sources checked often· verified 2026-09-21· EN

Training data for language models: filtering and dedup

Be able to build a filtered, deduplicated text corpus and measure its quality.

Prerequisites

Intuition

Raw web text is for the most part unusable: boilerplate, navigation menus, auto-generated text, spam and duplicates. The filtering is what decides the model's quality — more than the architecture does.

The pipeline, in order:

StepTypically discards
1. Extract the text from the HTMLtags, scripts, the menu
2. Language identificationthe wrong language
3. Quality heuristicstoo short, too many symbols, word lists
4. Deduplication30–60 % of everything
5. Toxicity and PII filtersharmful content, personal data
6. A contamination checktest data that has leaked in

Step 4 is the most impactful. Lee et al. (2021) showed that deduplication both reduces memorisation and improves the model — fewer repetitions means the model does not learn the same text a hundred times.

A rule of thumb that holds: ten times less but cleaned data beats ten times more uncleaned.

Formal

Deduplication at three levels:

LevelFindsMethod
Exactidentical documentsa hash of the normalised text
Near-duplicatesthe same text with small changesMinHash + LSH
Substring levelrecurring passages in different documentsa suffix array

MinHash approximates the Jaccard similarity between shingle sets, and LSH makes it possible to find the candidates without comparing every pair — without it, dedup at billion scale is impossible.

Quality heuristics (from the Gopher and C4 work) — simple rules that remove a lot of rubbish:

RuleDiscards
< 50 or > 100 000 wordsfragments and dumps
A mean word length outside 3–10 characterscode, noise, tables
> 90 % of the lines start with a bulletnavigation
< 80 % of the words in a word listnoise and code litter
A symbol share > 10 %scripts, tables
No punctuationlists and menus

Model-based filtering goes one step further: train a classifier on «good» text (Wikipedia, books) against random web text and keep what resembles the former. Effective, but it homogenises the corpus and risks systematically filtering out dialects, minority languages and informal registers.

Contamination is the check most often missing. If the test set has happened to end up in the training data, every evaluation is meaningless. Check it with n-gram overlap between the corpus and every benchmark you intend to use, and report the result.

Personal data in the corpus requires an explicit position: detect and mask emails, phone numbers, national ID numbers and addresses. That is not enough to make the corpus GDPR-safe, but it removes the most obvious.

Measure the corpus, not just its size:

MetricTells you
Total and unique tokensthe volume and the variation
The share remaining after each filter stepwhere the data disappears
The perplexity under a reference modelrubbish in the tails
The domain distributionwhat the model will know
The language distributionparticularly important for Swedish
The contamination rate per benchmarkwhether the evaluation holds

Code

import hashlib, re, unicodedata
from collections import Counter

def normalise(text: str) -> str:
    t = unicodedata.normalize("NFKC", text).lower()
    return re.sub(r"\s+", " ", t).strip()

def exact_dedup(documents):
    seen, out = set(), []
    for d in documents:
        h = hashlib.sha256(normalise(d).encode()).hexdigest()
        if h not in seen:
            seen.add(h); out.append(d)
    return out

# MinHash + LSH for near-duplicates
def shingles(text, n=5):
    words = normalise(text).split()
    return {" ".join(words[i:i + n]) for i in range(max(len(words) - n + 1, 1))}

def minhash(s, count=128, seed=0):
    import random
    rng = random.Random(seed)
    seeds = [rng.getrandbits(64) for _ in range(count)]
    sig = []
    for f in seeds:
        sig.append(min(hash((sh, f)) & 0xFFFFFFFF for sh in s) if s else 0)
    return tuple(sig)

def lsh_bands(sig, bands=16):
    r = len(sig) // bands
    return [hash(sig[i * r:(i + 1) * r]) for i in range(bands)]

def near_dedup(documents, bands=16, threshold=0.8):
    buckets, keep = {}, []
    for i, d in enumerate(documents):
        s = shingles(d)
        sig = minhash(s)
        candidates = set()
        keys = lsh_bands(sig, bands)
        for key in keys:
            candidates |= buckets.get(key, set())
        duplicate = any(
            len(s & shingles(documents[j])) / max(len(s | shingles(documents[j])), 1) > threshold
            for j in candidates)
        if not duplicate:
            keep.append(d)
            for key in keys:
                buckets.setdefault(key, set()).add(i)
    return keep

# Quality heuristics
SYMBOLS = set("#<>{}[]|\\@^~`")

def quality_ok(text, wordlist=None):
    words = text.split()
    if not 50 <= len(words) <= 100_000:
        return False, "length"
    if not 3 <= sum(len(w) for w in words) / len(words) <= 10:
        return False, "mean word length"
    lines = [l for l in text.splitlines() if l.strip()]
    if lines and sum(l.lstrip().startswith(("•", "-", "*")) for l in lines) / len(lines) > 0.9:
        return False, "bullet list"
    if sum(c in SYMBOLS for c in text) / max(len(text), 1) > 0.10:
        return False, "symbols"
    if not re.search(r"[.!?]", text):
        return False, "no punctuation"
    if wordlist and sum(w.strip(".,!?").lower() in wordlist for w in words) / len(words) < 0.80:
        return False, "word list"
    return True, "ok"

# A contamination check against the benchmarks
def contamination(corpus_ngrams: set, benchmark_texts, n=13):
    hits = 0
    for t in benchmark_texts:
        words = normalise(t).split()
        grams = {" ".join(words[i:i + n]) for i in range(max(len(words) - n + 1, 1))}
        if grams & corpus_ngrams:
            hits += 1
    return {"contaminated": hits, "of": len(benchmark_texts),
            "share": round(hits / max(len(benchmark_texts), 1), 4)}

# Mask the personal data
PATTERNS = {
    "EMAIL": r"\b[\w.+-]+@[\w-]+\.[\w.]+\b",
    "PHONE": r"\b0[\d\s-]{7,12}\b",
    "NATIONAL_ID": r"\b(19|20)?\d{6}[-+]?\d{4}\b",
}

def mask_pii(text):
    counts = Counter()
    for name, m in PATTERNS.items():
        text, n = re.subn(m, f"<{name}>", text)
        counts[name] += n
    return text, dict(counts)

# Report where the data disappears
def pipeline_report(documents, wordlist=None):
    steps = {"in": len(documents)}
    d = exact_dedup(documents); steps["after exact dedup"] = len(d)
    d = [x for x in d if quality_ok(x, wordlist)[0]]; steps["after quality"] = len(d)
    d = near_dedup(d); steps["after near-dedup"] = len(d)
    steps["share_remaining"] = round(len(d) / max(len(documents), 1), 3)
    return d, steps

Mastery means

  • Builds a filtering pipeline
  • Deduplicates at several levels
  • Measures corpus quality and contamination

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences