Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Language models — training and generation

Be able to explain next-token training, perplexity, sampling (temperature, top-p) and the difference between a base model and an instruction model.

Prerequisites

Intuition

A language model is a function from a token sequence to a probability distribution over the next token. Training: show it enormous amounts of text and punish it (cross-entropy) every time the correct next token got a low probability. No answer key is written by humans — the text is the answer key. That is called self-supervised training.

Perplexity = exp(mean cross-entropy): «how many equally likely alternatives the model is on average hesitating between». 20 is good on English, 2 is suspicious (memorisation or leaked test data).

Generation: draw a token from the distribution, append it, repeat. Temperature scales the logits before the softmax: T→0 becomes greedy (always the most likely), T = 1 is the model's own distribution, T > 1 is flatter and more creative. Top-p: draw only from the smallest set of tokens whose probabilities sum to p — cutting away the tail of nonsense.

A base model continues text. An instruction model has then been trained on (question, good answer) pairs and often preference data (RLHF/DPO) — it answers instead of continuing.

Code

import numpy as np

def softmax(z):
    z = z - z.max(); e = np.exp(z); return e / e.sum()

logits = np.array([3.0, 2.0, 0.5, -1.0])          # next-token logits for 4 tokens
for T in (0.2, 1.0, 2.0):
    print(T, np.round(softmax(logits / T), 3))
# 0.2 [0.993 0.007 0.    0.   ]  ← almost greedy
# 1.0 [0.68  0.25  0.056 0.013]
# 2.0 [0.48  0.29  0.14  0.065] ← flatter

def top_p(p, thr=0.9):
    idx = np.argsort(p)[::-1]; cum = np.cumsum(p[idx])
    keep = idx[: int(np.searchsorted(cum, thr)) + 1]
    q = np.zeros_like(p); q[keep] = p[keep]; return q / q.sum()

print(np.round(top_p(softmax(logits)), 3))       # [0.731 0.269 0. 0.] — the tail is gone

# perplexity from the token probabilities for the correct token:
p_correct = np.array([0.4, 0.1, 0.6, 0.05])
print(np.exp(-np.mean(np.log(p_correct))))       # ≈ 5.3

Formal

The model parameterises pθ(xt∣x<t)p_\theta(x_t \mid x_{<t}). The training objective is L(θ)=−∑tlog⁡pθ(xt∣x<t)\mathcal L(\theta) = -\sum_t \log p_\theta(x_t\mid x_{<t}) over the corpus; the probability of a whole sequence is the chain rule ∏tpθ(xt∣x<t)\prod_t p_\theta(x_t\mid x_{<t}). Perplexity on a test text of NN tokens: PPL=exp⁡ ⁣(1N∑t−log⁡pθ(xt∣x<t))\text{PPL} = \exp\!\big(\tfrac{1}{N}\sum_t -\log p_\theta(x_t\mid x_{<t})\big) — it is tokenizer-dependent, so only compare models with the same tokenizer. Temperature: pi∝exp⁡(zi/T)p_i \propto \exp(z_i/T). The scaling laws (Kaplan 2020, Hoffmann 2022) show that the test loss falls as a power law in parameters, data and compute, with an optimal data/parameter ratio (~20 tokens per parameter in Chinchilla).

Mastery means

  • Explains next-token training and perplexity
  • Describes temperature and top-p and their effects
  • Distinguishes a base model from an instruction model

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences