Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Base against instruction models

Be able to explain the difference between a pretrained base and an instruction-tuned model, and chat templates.

Prerequisites

Intuition

The base model is pretrained to predict the next token in web text. Give it «What is the capital of Sweden?» and it may answer with an answer — or with more questions, because that is often what texts with questions look like.

The instruction model has then been trained in two steps:

  1. SFT on (instruction, good answer) pairs → it learns to answer.
  2. Preference training (RLHF or DPO) on pairs where people ranked two answers → it learns which kind of answer is preferred: helpful, safe, well formatted.

Base models are used for further training and research. For everything else you choose the instruction variant (-instruct, -chat).

Code

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
msgs = [{"role": "system", "content": "You are a helpful tutor."},
        {"role": "user", "content": "What is a gradient?"}]
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
print(prompt)
# <|im_start|>system
# You are a helpful tutor.<|im_end|>
# <|im_start|>user
# What is a gradient?<|im_end|>
# <|im_start|>assistant

The chat template is not cosmetic. Every model family has its own (ChatML, Llama-3, Mistral …), and the model is trained on exactly those special tokens. Use the wrong template — or none at all — and the model loses measurably in quality and can start writing the user's lines for you. Always use apply_chat_template instead of building the string yourself.

Side effects of preference training worth knowing about: models become more verbose than necessary (length correlates with preference in the training data), more agreeable (sycophancy — agreeing with the user even when the user is wrong), and sometimes unnecessarily cautious. Those are trade-offs that have been made, not bugs.

Mastery means

  • Tells a base model from an instruction model
  • Uses the right chat template
  • Knows what RLHF/DPO add

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences