Dataset design for fine-tuning
Be able to build, clean and balance an instruction dataset, and measure how data quality affects the result.
Prerequisites
- BTraining data, features and labelsrequired
- EFine-tuning language modelsrequired
Intuition
In fine-tuning, data quality matters more than data quantity. The LIMA paper reached strong instruction following with 1 000 carefully chosen examples. Thousands of sloppy examples make the model worse, not better.
What «quality» means concretely:
- The answers are in exactly the form you want in production (length, tone, structure).
- No contradictions between examples.
- No errors — the model learns the errors verbatim.
- Coverage of the cases that actually occur, including the hard ones.
- Examples of declining when the material is not there.
The last one is nearly always forgotten: a model that has never seen an «I don't know» example will never answer that way.
Code
import hashlib, re
from collections import Counter
def normalise(t):
return re.sub(r"\s+", " ", t.lower()).strip()
def clean(examples, eval_texts, min_words=5, max_words=800):
seen, eval_h = set(), {hashlib.sha256(normalise(t).encode()).hexdigest() for t in eval_texts}
out, dropped = [], Counter()
for e in examples:
h = hashlib.sha256(normalise(e["instruction"] + e["answer"]).encode()).hexdigest()
n = len(e["answer"].split())
if h in seen: dropped["duplicate"] += 1; continue
if h in eval_h: dropped["leakage_vs_eval"] += 1; continue
if not (min_words <= n <= max_words): dropped["length"] += 1; continue
if "as a language model" in e["answer"].lower(): dropped["boilerplate"] += 1; continue
seen.add(h); out.append(e)
return out, dict(dropped)
cleaned, stats = clean(raw, eval_texts=[c["input"] for c in eval_cases])
print(len(raw), "→", len(cleaned), stats)
# 12400 → 9180 {'duplicate': 2010, 'leakage_vs_eval': 37, 'length': 980, 'boilerplate': 193}
Balancing: count the examples per category and per answer length. If 70 % of the examples are of one type, the model learns that type at the others' expense. Downsample the large category rather than upsampling the small ones (duplicates are learnt verbatim).
Measure that it helped: train on 25 %, 50 % and 100 % of the data and plot the curve. If it levels off, more data is not the answer — better data or a different method is. It is a two-hour experiment that often saves weeks.
Mastery means
- Builds and balances an instruction dataset
- Cleans out duplicates and leakage against the eval suite
- Measures how data quality affects the result
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — LIMA: Less Is More for Alignment — arXiv (open access; licence per article)
- arXiv — Training language models to follow instructions with human feedback — arXiv (open access; licence per article)