Skip to content
AI-grafen
EUniversityModel training and fine-tuning· about 60 min· fast-moving, sources checked often· verified 2026-09-20· EN

Fine-tuning language models

Be able to fine-tune a small model on instruction data, choose the hyperparameters and measure the improvement against the base model.

Prerequisites

Intuition

Fine-tuning = keep training a pretrained model on your data. The model already knows the language; you teach it format, domain and behaviour.

Instruction fine-tuning (SFT): the data is (instruction, answer) pairs. The loss is usually computed on the answer tokens only. The model learns to answer in your style. 500–5 000 good examples make more difference than 50 000 bad ones.

The hyperparameters that matter: the learning rate (1e-5 to 2e-4 depending on full/LoRA), 1–3 epochs (more → memorisation), an effective batch of 16–64, and a prompt template that is identical in training and at inference.

Always measure: the same eval before and after. No improvement on the eval = the fine-tune did nothing (or damaged something else — check a general eval too).

Code

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

name = "Qwen/Qwen2.5-0.5B"
tok = AutoTokenizer.from_pretrained(name); m = AutoModelForCausalLM.from_pretrained(name)
TEMPLATE = "### Instruction:\n{q}\n### Answer:\n"

def encode(q, a):
    p = tok(TEMPLATE.format(q=q))["input_ids"]; s = tok(a + tok.eos_token)["input_ids"]
    ids = torch.tensor(p + s); labels = ids.clone(); labels[: len(p)] = -100   # no loss on the prompt
    return ids, labels

opt = torch.optim.AdamW(m.parameters(), lr=1e-5)
for epoch in range(2):
    for q, a in data:
        ids, labels = encode(q, a)
        loss = m(input_ids=ids[None], labels=labels[None]).loss
        loss.backward(); opt.step(); opt.zero_grad()

# inference — the SAME template
x = tok(TEMPLATE.format(q="What is a tensor?"), return_tensors="pt")
print(tok.decode(m.generate(**x, max_new_tokens=60)[0][x["input_ids"].shape[1]:]))

In practice: LoRA (the next level) for memory, gradient accumulation for the batch, bf16, and an eval harness that is run before and after.

Mastery means

  • Fine-tunes a small model on instruction data
  • Chooses the lr and the number of epochs and formats the data with a prompt template
  • Measures the improvement against the base model with an eval

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences