How a language model is trained — at upper-secondary level
Be able to explain next-word training, why the amount of text matters, and what fine-tuning is.
Prerequisites
Intuition
A language model is trained on a single task: guess the next word.
"The capital of Sweden is ___"
The model gives a probability to every possible next word. When it guesses wrong the weights are nudged a little. Repeat a few trillion times.
What is surprising is what comes for free. To get good at guessing the next word, the model has to learn:
| To guess right here | it has to know |
|---|---|
| «The capital of Sweden is …» | geography |
| «2 + 2 = …» | arithmetic |
| «He said that she …» | grammar and reference |
| «def add(a, b): return …» | programming |
| «The patient had a fever and …» | medical terminology |
Nobody taught it any of that. It follows from the next word being easier to guess if you understand what the text is about.
No database. The model does not look the answer up — it has no stored text. The knowledge sits in billions of weights, roughly as you remember that Stockholm is the capital without having a list in your head. And just like your memory, it can mix things up.
Formal
Three phases, in order:
| Phase | Data | Cost | Gives |
|---|---|---|---|
| Pretraining | trillions of words from the web, books, code | months, millions of kronor | language and knowledge |
| Instruction tuning | tens of thousands of (question, good answer) pairs | days | follows instructions instead of continuing the text |
| Preference tuning | humans rank the answers | days | helpfulness, tone, safety |
After pretraining alone the model is a text continuer. Ask it «What is the capital of Sweden?» and it may answer «What is the capital of Norway? What is Denmark's …» — it continues the pattern instead of answering. It is instruction tuning that turns it into an assistant.
Why scale matters. The quality follows roughly a power law in three quantities: the number of parameters, the amount of training data and the compute. Ten times the resources gives a smooth, predictable improvement — but no abrupt win, and the returns diminish.
For a long time the belief was that only the parameters counted. The Chinchilla result (2022) showed that the amount of data was at least as important, and that most models were undertrained on data. That changed how models have been built ever since: more tokens, not just more weights.
Three limitations that follow directly from how the training works:
| Limitation | Why |
|---|---|
| Knowledge cut-off | the model knows nothing after training ended |
| Hallucinations | it guesses the most likely text, not the truth — a plausible invention is exactly what the task rewards |
| Bias | it reflects the text it was trained on, including that text's slant |
The middle one is worth pausing on: hallucinations are not a bug to be fixed away, but a direct consequence of the task being «produce likely text». Counteracting them requires something outside the model — retrieved sources, tools, or a human who checks.
Interactive
Four experiments in any chatbot. They show how the training works, from the inside.
1. The knowledge cut-off. Ask about something that happened in the last few months. The answer reveals where the training data ends — or that the model was allowed to search the web.
2. Hallucination on demand. Ask about something invented: «Tell me about Södertälje University's research on quantum biology.» Many models produce a well-formulated answer about something that does not exist. Then ask «are you sure?» and see whether the answer changes.
3. Probability, not truth. Ask for «ten English words that end in -ing». Check each one. Mistakes often turn up — the words sound right, which is exactly what the model optimises for.
4. The text continuer under the surface. Write an incomplete sentence with no question: «Once upon a time there was a» and see what happens. Then write the same thing but start with «Do not continue, answer this instead:».
What the experiments show together: the model is not a reference work that is sometimes wrong. It is a machine that produces likely text, and which is therefore right when the likely thing is also true. For everything else you need sources.
Write down one example from each experiment — they are more convincing than any explanation when you have to explain to someone else why you cannot trust a chatbot blindly.
Mastery means
- Explains next-token training
- Describes the three training phases
- Explains why the amount of data and the scale matter
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Training Compute-Optimal Large Language Models (Chinchilla) — arXiv (open access; licence per article)
- Skolverket — About AI in school (in Swedish) — Skolverket's open terms
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0