Fine-tuning — an overview for upper secondary
Be able to explain why you fine-tune instead of retraining, and what it requires.
Prerequisites
Intuition
Training a language model from scratch costs millions and takes months. Fine-tuning starts from a finished model and adjusts it a little — hours and a few hundred kronor.
The analogy: pre-training is all your schooling, fine-tuning is a week's induction at a new job.
Three ways of getting a model to do what you want, in order of cost:
| The method | The cost | Good at | Bad at |
|---|---|---|---|
| A prompt | free | quick adjustments, instructions | consistent style, large changes |
| RAG | low | current and your own facts | style, format, tone |
| Fine-tuning | medium | style, format, tone, domain language | adding facts |
The most common misconception is that fine-tuning is the way to teach the model new facts. It works badly: facts in fine-tuning data are memorised unreliably, get mixed up and become impossible to update. If you want the model to know your product data — use RAG.
The rule: fine-tune for how the model should sound and behave. Retrieve what has to be true and current.
Formal
What a successful fine-tuning requires:
| The part | A guide value |
|---|---|
| Examples | 500–5 000 of high quality |
| The quality | more important than the quantity — 500 reviewed beat 50 000 sloppy ones |
| The format | (instruction, answer) pairs in the model's chat template |
| The evaluation | a test set that has never been used in the training |
| The computation | one GPU for a few hours for a small model with LoRA |
LoRA makes it affordable: instead of updating all the billions of weights, small additions (a few million parameters) are trained alongside. The result is nearly as good, the memory requirement falls drastically, and you can have several fine-tunings of the same base model and switch between them.
Three things that usually go wrong:
- Too few examples. Below a couple of hundred nothing measurable usually happens.
- Too many epochs. The model memorises the training examples and becomes worse at everything else.
- No baseline. The most common discovery after a fine-tuning is that a good prompt would have given the same result. Always measure against the base model with few-shot before drawing conclusions.
Catastrophic forgetting is the fourth and most insidious: the model becomes better at your task and worse at things you are not measuring. So the evaluation should always have two parts — the target task and a durability suite of general tasks.
The decision order in practice: try a prompt first. If that is not enough and the problem is about facts — RAG. If it is about style, format or domain language — fine-tuning. Often both RAG and fine-tuning are needed, and they then solve different things.
Interactive
Decide the right method for five real cases. Read, choose, then compare.
| # | The situation | Prompt / RAG / Fine-tuning? |
|---|---|---|
| 1 | The model should answer questions about your school's rules | ? |
| 2 | The answers should always be at most three sentences and address the reader directly | ? |
| 3 | The model should know your 400 products' prices, which change every week | ? |
| 4 | The model should write in your authority's established style | ? |
| 5 | The model should handle Swedish legal terminology it often gets wrong | ? |
The answers and the reasoning:
- RAG — the rules are facts that exist in a document, and they should be updatable without retraining.
- A prompt — an instruction is plenty. Fine-tuning would be shooting a mosquito with a cannon.
- RAG, without a doubt — prices that change every week can hardly live in the weights.
- Fine-tuning — style is exactly what fine-tuning is good at, and hard to describe exhaustively in a prompt.
- Both — fine-tune on the terminology so that the model uses the right words, and use RAG against the statute text so that the content is correct and current.
The pattern: if you are asking «what should the model know?» the answer is RAG. If you are asking «how should it sound?» the answer is fine-tuning. If it is only a small adjustment of the behaviour — a prompt.
Mastery means
- Explains the difference between pre-training and fine-tuning
- Knows what fine-tuning requires and does not solve
- Chooses between a prompt, RAG and fine-tuning
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — LoRA: Low-Rank Adaptation of Large Language Models — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0
- Skolverket — About AI in school (in Swedish) — Skolverket's open terms