Tokenisation
Be able to explain how text is split into tokens (BPE), why the vocabulary size matters, and count the tokens in a text.
Prerequisites
Intuition
A language model does not see letters or words — it sees tokens: pieces of text from a fixed vocabulary (30 000–200 000 entries). «cats» may be one token, «kittens» three: «kit», «ten», «s».
BPE (byte-pair encoding) builds the vocabulary: start with individual characters, merge the most common pair into a new token, repeat thousands of times. Common words become one token, uncommon ones are built from parts. Nothing is «unknown».
The trade-off: a larger vocabulary → shorter sequences (cheaper) but more parameters in the embedding and rare tokens that barely get trained. Swedish is often tokenised into more pieces than English — that costs context and money.
Code
# Minimal BPE training (illustrative)
from collections import Counter
def bpe_train(word_freq, steps):
vocab = {tuple(w) + ("</w>",): f for w, f in word_freq.items()}
merges = []
for _ in range(steps):
pairs = Counter()
for w, f in vocab.items():
for a, b in zip(w, w[1:]):
pairs[(a, b)] += f
if not pairs: break
(a, b), _ = pairs.most_common(1)[0]
merges.append((a, b))
vocab = {tuple(" ".join(w).replace(f"{a} {b}", a + b).split()): f for w, f in vocab.items()}
return merges
print(bpe_train({"low": 5, "lower": 2, "lowest": 3}, 3))
# [('l', 'o'), ('lo', 'w'), ('low', '</w>')]
With a real tokenizer: from transformers import AutoTokenizer; tok = AutoTokenizer.from_pretrained("gpt2"); len(tok("Hello world")["input_ids"]).
Mastery means
- Explains how BPE builds a vocabulary of sub-words
- Counts the tokens in a text with a tokenizer
- Explains why the vocabulary size is a trade-off
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Neural Machine Translation of Rare Words with Subword Units — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0
Leads to
Part of the goals (22)
- Build a transformer from scratch
- Language models in practice
- Multimodal systems
- Fine-tune and run your own models
- Build a RAG system you can trust
- AI safety in practice
- Fine-tune a model with LoRA
- Responsible AI in practice
- Evals in practice
- Build an agent you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- An AI service in operation
- Run models more cheaply: quantisation
- Build a memory system for an agent
- Build an AI service that survives production
- Deep reinforcement learning
- Frontier Lab — an independent research project
- Reproduce a paper
- Interpreting a language model
- Statistics for experiments