Skip to content
AI-grafen
FAI engineeringData handling· about 90 min· evolving, reviewed regularly· verified 2026-09-20· EN

Dataset design for fine-tuning

Be able to build, clean and balance an instruction dataset, and measure how data quality affects the result.

Prerequisites

Intuition

In fine-tuning, data quality matters more than data quantity. The LIMA paper reached strong instruction following with 1 000 carefully chosen examples. Thousands of sloppy examples make the model worse, not better.

What «quality» means concretely:

  • The answers are in exactly the form you want in production (length, tone, structure).
  • No contradictions between examples.
  • No errors — the model learns the errors verbatim.
  • Coverage of the cases that actually occur, including the hard ones.
  • Examples of declining when the material is not there.

The last one is nearly always forgotten: a model that has never seen an «I don't know» example will never answer that way.

Code

import hashlib, re
from collections import Counter

def normalise(t):
    return re.sub(r"\s+", " ", t.lower()).strip()

def clean(examples, eval_texts, min_words=5, max_words=800):
    seen, eval_h = set(), {hashlib.sha256(normalise(t).encode()).hexdigest() for t in eval_texts}
    out, dropped = [], Counter()
    for e in examples:
        h = hashlib.sha256(normalise(e["instruction"] + e["answer"]).encode()).hexdigest()
        n = len(e["answer"].split())
        if h in seen:                        dropped["duplicate"] += 1; continue
        if h in eval_h:                      dropped["leakage_vs_eval"] += 1; continue
        if not (min_words <= n <= max_words): dropped["length"] += 1; continue
        if "as a language model" in e["answer"].lower(): dropped["boilerplate"] += 1; continue
        seen.add(h); out.append(e)
    return out, dict(dropped)

cleaned, stats = clean(raw, eval_texts=[c["input"] for c in eval_cases])
print(len(raw), "→", len(cleaned), stats)
# 12400 → 9180 {'duplicate': 2010, 'leakage_vs_eval': 37, 'length': 980, 'boilerplate': 193}

Balancing: count the examples per category and per answer length. If 70 % of the examples are of one type, the model learns that type at the others' expense. Downsample the large category rather than upsampling the small ones (duplicates are learnt verbatim).

Measure that it helped: train on 25 %, 50 % and 100 % of the data and plot the curve. If it levels off, more data is not the answer — better data or a different method is. It is a two-hour experiment that often saves weeks.

Mastery means

  • Builds and balances an instruction dataset
  • Cleans out duplicates and leakage against the eval suite
  • Measures how data quality affects the result

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences