Base against instruction models
Be able to explain the difference between a pretrained base and an instruction-tuned model, and chat templates.
Prerequisites
Intuition
The base model is pretrained to predict the next token in web text. Give it «What is the capital of Sweden?» and it may answer with an answer — or with more questions, because that is often what texts with questions look like.
The instruction model has then been trained in two steps:
- SFT on (instruction, good answer) pairs → it learns to answer.
- Preference training (RLHF or DPO) on pairs where people ranked two answers → it learns which kind of answer is preferred: helpful, safe, well formatted.
Base models are used for further training and research. For everything else you choose the instruction variant (-instruct, -chat).
Code
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("Qwen/Qwen2.5-7B-Instruct")
msgs = [{"role": "system", "content": "You are a helpful tutor."},
{"role": "user", "content": "What is a gradient?"}]
prompt = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True)
print(prompt)
# <|im_start|>system
# You are a helpful tutor.<|im_end|>
# <|im_start|>user
# What is a gradient?<|im_end|>
# <|im_start|>assistant
The chat template is not cosmetic. Every model family has its own (ChatML, Llama-3, Mistral …), and the model is trained on exactly those special tokens. Use the wrong template — or none at all — and the model loses measurably in quality and can start writing the user's lines for you. Always use apply_chat_template instead of building the string yourself.
Side effects of preference training worth knowing about: models become more verbose than necessary (length correlates with preference in the training data), more agreeable (sycophancy — agreeing with the user even when the user is wrong), and sometimes unnecessarily cautious. Those are trade-offs that have been made, not bugs.
Mastery means
- Tells a base model from an instruction model
- Uses the right chat template
- Knows what RLHF/DPO add
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Training language models to follow instructions with human feedback — arXiv (open access; licence per article)
- arXiv — Direct Preference Optimization — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0