Skip to content
AI-grafen
FAI engineeringModel training and fine-tuning· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Pre-training in practice

Be able to describe the data mixture, the scaling laws and the compute budget for pre-training a small language model.

Prerequisites

Intuition

Pre-training is teaching a model language from scratch — next-token prediction on hundreds of billions of tokens. It is a different activity from fine-tuning: months of GPU time, a data team, and decisions that cannot be undone afterwards.

The three decisions that decide the result:

  1. The data mixture. Not «all the text we can find» but weighted shares: web text, code, books, science, multilingual. Code in the mixture improves reasoning even for models that are not going to write code.
  2. The compute budget and how it is divided between the model size and the data volume.
  3. The data quality — deduplication, filtering, and keeping the test data out.

Formal

Scaling laws. Kaplan et al. (2020) showed that the test loss follows a power law in the parameters N, the data D and the computation C. Hoffmann et al. (2022, «Chinchilla») corrected the allocation: for a given budget C≈6NDC \approx 6ND, N and D should be scaled proportionally — roughly 20 tokens per parameter.

That means GPT-3 (175 B parameters, 300 B tokens) was heavily undertrained; a 70 B model on 1.4 T tokens performed better at the same compute cost.

Today models are often trained far beyond the Chinchilla optimum (Llama 3: 8 B parameters on 15 T tokens ≈ 1 875 tokens/parameter). The reason is inference economics: a smaller model trained longer costs less per call for ever after, and that cost dominates the total for a model that is used a lot.

Data processing in order of magnitude:

The stepThe effect
Language identification and filtering−50–70 % of the raw web data
Deduplication (exact + fuzzy/MinHash)−20–30 %, a measurable improvement
Quality filtering (a classifier, heuristics)−30–50 %
Removing benchmark overlapsmall in volume, decisive for credibility
PII removala requirement, not an optimisation

FineWeb and Dolma show that the filtering often makes a larger difference than the choice of architecture.

Code

def chinchilla(budget_flops):
    """C ≈ 6ND and D ≈ 20N → N = sqrt(C / 120)"""
    N = (budget_flops / 120) ** 0.5
    return {"parameters_B": round(N / 1e9, 2), "tokens_B": round(20 * N / 1e9, 1)}

for hours, gpus in ((1_000, 8), (10_000, 64)):
    flops = hours * 3600 * gpus * 300e12 * 0.4        # an H100 at ~300 TFLOPs bf16, 40 % MFU
    print(hours, gpus, chinchilla(flops))
# 1000 8   {'parameters_B': 0.29, 'tokens_B': 5.9}
# 10000 64 {'parameters_B': 2.63, 'tokens_B': 52.5}

# An example data mixture (weighted shares, not raw sizes)
MIXTURE = {"web_filtered": 0.55, "code": 0.15, "books": 0.12,
           "science": 0.08, "multilingual": 0.08, "conversation": 0.02}
assert abs(sum(MIXTURE.values()) - 1.0) < 1e-9

The arithmetic shows why few people train their own base models: 10 000 GPU hours is enough for a 2.6 B model — considerably worse than open models you can download for free. Pre-training pays off only with unique data or unique requirements.

Mastery means

  • Describes the data mixture and the compute budget for pre-training
  • Uses scaling laws to allocate the budget
  • Knows what separates pre-training from fine-tuning

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences