Skip to content
AI-grafen
DAI developerLanguage models· about 45 min· evolving, reviewed regularly· verified 2026-09-20· EN

Tokenisation

Be able to explain how text is split into tokens (BPE), why the vocabulary size matters, and count the tokens in a text.

Prerequisites

Intuition

A language model does not see letters or words — it sees tokens: pieces of text from a fixed vocabulary (30 000–200 000 entries). «cats» may be one token, «kittens» three: «kit», «ten», «s».

BPE (byte-pair encoding) builds the vocabulary: start with individual characters, merge the most common pair into a new token, repeat thousands of times. Common words become one token, uncommon ones are built from parts. Nothing is «unknown».

The trade-off: a larger vocabulary → shorter sequences (cheaper) but more parameters in the embedding and rare tokens that barely get trained. Swedish is often tokenised into more pieces than English — that costs context and money.

Code

# Minimal BPE training (illustrative)
from collections import Counter

def bpe_train(word_freq, steps):
    vocab = {tuple(w) + ("</w>",): f for w, f in word_freq.items()}
    merges = []
    for _ in range(steps):
        pairs = Counter()
        for w, f in vocab.items():
            for a, b in zip(w, w[1:]):
                pairs[(a, b)] += f
        if not pairs: break
        (a, b), _ = pairs.most_common(1)[0]
        merges.append((a, b))
        vocab = {tuple(" ".join(w).replace(f"{a} {b}", a + b).split()): f for w, f in vocab.items()}
    return merges

print(bpe_train({"low": 5, "lower": 2, "lowest": 3}, 3))
# [('l', 'o'), ('lo', 'w'), ('low', '</w>')]

With a real tokenizer: from transformers import AutoTokenizer; tok = AutoTokenizer.from_pretrained("gpt2"); len(tok("Hello world")["input_ids"]).

Mastery means

  • Explains how BPE builds a vocabulary of sub-words
  • Counts the tokens in a text with a tokenizer
  • Explains why the vocabulary size is a trade-off

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences