Language models — training and generation
Be able to explain next-token training, perplexity, sampling (temperature, top-p) and the difference between a base model and an instruction model.
Prerequisites
- DLoss functionsrequired
- DTokenisationrequired
- DTransformers — the architecturerequired
Intuition
A language model is a function from a token sequence to a probability distribution over the next token. Training: show it enormous amounts of text and punish it (cross-entropy) every time the correct next token got a low probability. No answer key is written by humans — the text is the answer key. That is called self-supervised training.
Perplexity = exp(mean cross-entropy): «how many equally likely alternatives the model is on average hesitating between». 20 is good on English, 2 is suspicious (memorisation or leaked test data).
Generation: draw a token from the distribution, append it, repeat. Temperature scales the logits before the softmax: T→0 becomes greedy (always the most likely), T = 1 is the model's own distribution, T > 1 is flatter and more creative. Top-p: draw only from the smallest set of tokens whose probabilities sum to p — cutting away the tail of nonsense.
A base model continues text. An instruction model has then been trained on (question, good answer) pairs and often preference data (RLHF/DPO) — it answers instead of continuing.
Code
import numpy as np
def softmax(z):
z = z - z.max(); e = np.exp(z); return e / e.sum()
logits = np.array([3.0, 2.0, 0.5, -1.0]) # next-token logits for 4 tokens
for T in (0.2, 1.0, 2.0):
print(T, np.round(softmax(logits / T), 3))
# 0.2 [0.993 0.007 0. 0. ] ← almost greedy
# 1.0 [0.68 0.25 0.056 0.013]
# 2.0 [0.48 0.29 0.14 0.065] ← flatter
def top_p(p, thr=0.9):
idx = np.argsort(p)[::-1]; cum = np.cumsum(p[idx])
keep = idx[: int(np.searchsorted(cum, thr)) + 1]
q = np.zeros_like(p); q[keep] = p[keep]; return q / q.sum()
print(np.round(top_p(softmax(logits)), 3)) # [0.731 0.269 0. 0.] — the tail is gone
# perplexity from the token probabilities for the correct token:
p_correct = np.array([0.4, 0.1, 0.6, 0.05])
print(np.exp(-np.mean(np.log(p_correct)))) # ≈ 5.3
Formal
The model parameterises . The training objective is over the corpus; the probability of a whole sequence is the chain rule . Perplexity on a test text of tokens: — it is tokenizer-dependent, so only compare models with the same tokenizer. Temperature: . The scaling laws (Kaplan 2020, Hoffmann 2022) show that the test loss falls as a power law in parameters, data and compute, with an optimal data/parameter ratio (~20 tokens per parameter in Chinchilla).
Mastery means
- Explains next-token training and perplexity
- Describes temperature and top-p and their effects
- Distinguishes a base model from an instruction model
Sign in to do the exercises and build your mastery up.
Sources
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- arXiv — Language Models are Few-Shot Learners — arXiv (open access; licence per article)
- arXiv — Training Compute-Optimal Large Language Models — arXiv (open access; licence per article)
Leads to
- EBase against instruction models
- EFine-tuning language models
- EMultilinguality and Swedish models
- EHallucinations — causes and countermeasures
- EModel evaluation
- EPerplexity
- ERAG — retrieval-augmented generation
- ESampling: temperature, top-k, top-p, beam
- ELanguage models for code
- ETool use
- FImage description and visual questions
- FPre-training in practice
Part of the goals (21)
- Language models in practice
- Multimodal systems
- Fine-tune and run your own models
- Build a RAG system you can trust
- AI safety in practice
- Fine-tune a model with LoRA
- Responsible AI in practice
- Evals in practice
- Build an agent you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- An AI service in operation
- Run models more cheaply: quantisation
- Build a memory system for an agent
- Build an AI service that survives production
- Deep reinforcement learning
- Frontier Lab — an independent research project
- Reproduce a paper
- Interpreting a language model
- Statistics for experiments