Skip to content
AI-grafen
FAI engineeringData handling· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Synthetic data

Be able to generate and filter synthetic training data with a model, and detect model collapse.

Prerequisites

Intuition

Synthetic data is examples generated by a model instead of collected from reality. It is used to cover cases that are missing, to scale small datasets up, and to avoid personal data.

Where it works well: when there is a verifiable answer key. Mathematics (check the answer), code (run the tests), structured extraction (validate against the schema). Then you can generate a lot and filter hard — and the filter is what creates the value.

Where it works badly: open text with no key. Then the data inherits the model's own faults and stylistic quirks, and reinforces them.

The basic rule: generate broadly, filter hard, keep little. A dataset where 90 % has been thrown away is often better than one where everything was kept.

Formal

Model collapse (Shumailov et al. 2024): if you train repeated generations of models on the previous generation's output, the tails of the distribution are lost first — rare but valid patterns — and after that the distribution gradually narrows. The model becomes ever more generic and loses variation.

The mechanism is statistical: every generation samples from an estimated distribution, and the sampling errors accumulate. The effect appears after just a few generations in pure synthetic loops.

The countermeasures:

  • Always mix real data in (the original corpus is kept).
  • Verify the synthetic examples against an external key where one exists.
  • Measure the diversity of the generated data — n-gram entropy, embedding spread, the share of unique answers.
  • Use a stronger model as the generator than the one being trained (distillation), not the same model in a loop.

Licences and copyright: many providers' terms forbid using their output to train competing models. Check the terms before a synthetic dataset is built — it is a common and costly oversight.

Code

import numpy as np
from collections import Counter

def generate_and_filter(llm, seed_examples, n=5000, verify=None):
    candidates = [llm.generate_variant(rng_choice(seed_examples)) for _ in range(n)]
    kept, discarded = [], Counter()
    seen = set()
    for c in candidates:
        key = normalise(c["instruction"])
        if key in seen:                        discarded["duplicate"] += 1; continue
        if verify and not verify(c):           discarded["verification"] += 1; continue
        if len(c["answer"].split()) < 5:       discarded["too_short"] += 1; continue
        seen.add(key); kept.append(c)
    return kept, dict(discarded)

def diversity(examples, n=3):
    """The share of unique n-grams — if it falls sharply against real data the distribution is too narrow."""
    grams = [tuple(t[i:i+n]) for t in (e["answer"].split() for e in examples)
             for i in range(max(len(t) - n + 1, 0))]
    return len(set(grams)) / max(len(grams), 1)

syn, stats = generate_and_filter(llm, seeds, n=5000, verify=run_tests)
print(len(syn), stats, round(diversity(syn), 3), round(diversity(real_data), 3))
# 1180 {'duplicate': 2310, 'verification': 1290, 'too_short': 220} 0.41 0.63
#                                          ↑ 76 % discarded   ↑ a lower diversity than the real data

The diversity comparison against real data is the simplest early warning that the synthetic data is too narrow.

Mastery means

  • Generates and filters synthetic training data
  • Recognises model collapse
  • Knows when synthetic data helps and when it harms

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences