Skip to content
AI-grafen
EUniversityModel training and fine-tuning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Instruction fine-tuning (SFT)

Be able to fine-tune a small model on instruction data and measure the improvement.

Prerequisites

Intuition

SFT (supervised fine-tuning) teaches the model to answer instead of continuing a text. The data is (instruction, answer) pairs, and the recipe is simple — but three details decide the result:

  1. Mask the prompt. The loss should only be computed on the answer tokens. Otherwise the model learns to generate questions just as readily as answers.
  2. Use the right chat template — the same in training and at inference. The wrong template costs measurably without showing up as an error.
  3. Few epochs. 1–3. More gives memorisation, and on 1 000 examples it shows already at epoch 4.

Quality beats quantity: 1 000 carefully reviewed pairs beat 50 000 sloppy ones.

Code

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

name = "Qwen/Qwen2.5-0.5B-Instruct"
tok = AutoTokenizer.from_pretrained(name)
m = AutoModelForCausalLM.from_pretrained(name, torch_dtype=torch.bfloat16, device_map="auto")

def encode(instruction: str, answer: str, max_len=1024):
    prompt = tok.apply_chat_template([{"role": "user", "content": instruction}],
                                     tokenize=False, add_generation_prompt=True)
    p_ids = tok(prompt, add_special_tokens=False)["input_ids"]
    a_ids = tok(answer + tok.eos_token, add_special_tokens=False)["input_ids"]
    ids = (p_ids + a_ids)[:max_len]
    labels = ids.copy()
    labels[:len(p_ids)] = [-100] * min(len(p_ids), len(labels))   # mask the prompt
    return torch.tensor(ids), torch.tensor(labels)

opt = torch.optim.AdamW(m.parameters(), lr=1e-5)
for epoch in range(2):
    for instr, answer in data:
        ids, labels = encode(instr, answer)
        loss = m(input_ids=ids[None].cuda(), labels=labels[None].cuda()).loss
        loss.backward()
        torch.nn.utils.clip_grad_norm_(m.parameters(), 1.0)
        opt.step(); opt.zero_grad()

The evaluation — three suites, always:

The suiteWhat it answers
The target eval (≥ 50 cases)did the model get better at the task?
The few-shot baselinewas the fine-tuning needed at all?
Durability (≥ 50 cases)did it lose anything else?

The middle one is nearly always forgotten — and is the one that sometimes shows that a good prompt would have been enough.

Mastery means

  • Fine-tunes a model on instruction data
  • Masks the prompt in the loss
  • Measures the improvement against a baseline and the durability

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences