Pre-training in practice
Be able to describe the data mixture, the scaling laws and the compute budget for pre-training a small language model.
Prerequisites
- EOptimisers: momentum, Adam, schedulingrequired
- ELanguage models — training and generationrequired
- FDistributed traininghelpful
Intuition
Pre-training is teaching a model language from scratch — next-token prediction on hundreds of billions of tokens. It is a different activity from fine-tuning: months of GPU time, a data team, and decisions that cannot be undone afterwards.
The three decisions that decide the result:
- The data mixture. Not «all the text we can find» but weighted shares: web text, code, books, science, multilingual. Code in the mixture improves reasoning even for models that are not going to write code.
- The compute budget and how it is divided between the model size and the data volume.
- The data quality — deduplication, filtering, and keeping the test data out.
Formal
Scaling laws. Kaplan et al. (2020) showed that the test loss follows a power law in the parameters N, the data D and the computation C. Hoffmann et al. (2022, «Chinchilla») corrected the allocation: for a given budget , N and D should be scaled proportionally — roughly 20 tokens per parameter.
That means GPT-3 (175 B parameters, 300 B tokens) was heavily undertrained; a 70 B model on 1.4 T tokens performed better at the same compute cost.
Today models are often trained far beyond the Chinchilla optimum (Llama 3: 8 B parameters on 15 T tokens ≈ 1 875 tokens/parameter). The reason is inference economics: a smaller model trained longer costs less per call for ever after, and that cost dominates the total for a model that is used a lot.
Data processing in order of magnitude:
| The step | The effect |
|---|---|
| Language identification and filtering | −50–70 % of the raw web data |
| Deduplication (exact + fuzzy/MinHash) | −20–30 %, a measurable improvement |
| Quality filtering (a classifier, heuristics) | −30–50 % |
| Removing benchmark overlap | small in volume, decisive for credibility |
| PII removal | a requirement, not an optimisation |
FineWeb and Dolma show that the filtering often makes a larger difference than the choice of architecture.
Code
def chinchilla(budget_flops):
"""C ≈ 6ND and D ≈ 20N → N = sqrt(C / 120)"""
N = (budget_flops / 120) ** 0.5
return {"parameters_B": round(N / 1e9, 2), "tokens_B": round(20 * N / 1e9, 1)}
for hours, gpus in ((1_000, 8), (10_000, 64)):
flops = hours * 3600 * gpus * 300e12 * 0.4 # an H100 at ~300 TFLOPs bf16, 40 % MFU
print(hours, gpus, chinchilla(flops))
# 1000 8 {'parameters_B': 0.29, 'tokens_B': 5.9}
# 10000 64 {'parameters_B': 2.63, 'tokens_B': 52.5}
# An example data mixture (weighted shares, not raw sizes)
MIXTURE = {"web_filtered": 0.55, "code": 0.15, "books": 0.12,
"science": 0.08, "multilingual": 0.08, "conversation": 0.02}
assert abs(sum(MIXTURE.values()) - 1.0) < 1e-9
The arithmetic shows why few people train their own base models: 10 000 GPU hours is enough for a 2.6 B model — considerably worse than open models you can download for free. Pre-training pays off only with unique data or unique requirements.
Mastery means
- Describes the data mixture and the compute budget for pre-training
- Uses scaling laws to allocate the budget
- Knows what separates pre-training from fine-tuning
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Training Compute-Optimal Large Language Models (Chinchilla) — arXiv (open access; licence per article)
- arXiv — Scaling Laws for Neural Language Models — arXiv (open access; licence per article)
- arXiv — The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale — arXiv (open access; licence per article)