Synthetic data
Be able to generate and filter synthetic training data with a model, and detect model collapse.
Prerequisites
- FDataset design for fine-tuningrequired
- FEvals for language models and agentsrequired
Intuition
Synthetic data is examples generated by a model instead of collected from reality. It is used to cover cases that are missing, to scale small datasets up, and to avoid personal data.
Where it works well: when there is a verifiable answer key. Mathematics (check the answer), code (run the tests), structured extraction (validate against the schema). Then you can generate a lot and filter hard — and the filter is what creates the value.
Where it works badly: open text with no key. Then the data inherits the model's own faults and stylistic quirks, and reinforces them.
The basic rule: generate broadly, filter hard, keep little. A dataset where 90 % has been thrown away is often better than one where everything was kept.
Formal
Model collapse (Shumailov et al. 2024): if you train repeated generations of models on the previous generation's output, the tails of the distribution are lost first — rare but valid patterns — and after that the distribution gradually narrows. The model becomes ever more generic and loses variation.
The mechanism is statistical: every generation samples from an estimated distribution, and the sampling errors accumulate. The effect appears after just a few generations in pure synthetic loops.
The countermeasures:
- Always mix real data in (the original corpus is kept).
- Verify the synthetic examples against an external key where one exists.
- Measure the diversity of the generated data — n-gram entropy, embedding spread, the share of unique answers.
- Use a stronger model as the generator than the one being trained (distillation), not the same model in a loop.
Licences and copyright: many providers' terms forbid using their output to train competing models. Check the terms before a synthetic dataset is built — it is a common and costly oversight.
Code
import numpy as np
from collections import Counter
def generate_and_filter(llm, seed_examples, n=5000, verify=None):
candidates = [llm.generate_variant(rng_choice(seed_examples)) for _ in range(n)]
kept, discarded = [], Counter()
seen = set()
for c in candidates:
key = normalise(c["instruction"])
if key in seen: discarded["duplicate"] += 1; continue
if verify and not verify(c): discarded["verification"] += 1; continue
if len(c["answer"].split()) < 5: discarded["too_short"] += 1; continue
seen.add(key); kept.append(c)
return kept, dict(discarded)
def diversity(examples, n=3):
"""The share of unique n-grams — if it falls sharply against real data the distribution is too narrow."""
grams = [tuple(t[i:i+n]) for t in (e["answer"].split() for e in examples)
for i in range(max(len(t) - n + 1, 0))]
return len(set(grams)) / max(len(grams), 1)
syn, stats = generate_and_filter(llm, seeds, n=5000, verify=run_tests)
print(len(syn), stats, round(diversity(syn), 3), round(diversity(real_data), 3))
# 1180 {'duplicate': 2310, 'verification': 1290, 'too_short': 220} 0.41 0.63
# ↑ 76 % discarded ↑ a lower diversity than the real data
The diversity comparison against real data is the simplest early warning that the synthetic data is too narrow.
Mastery means
- Generates and filters synthetic training data
- Recognises model collapse
- Knows when synthetic data helps and when it harms
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — AI models collapse when trained on recursively generated data — arXiv (open access; licence per article)
- arXiv — Self-Instruct: Aligning Language Models with Self-Generated Instructions — arXiv (open access; licence per article)