Skip to content
AI-grafen
EUniversityLanguage models· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Byte-pair encoding

Be able to implement BPE training and tokenisation and explain the vocabulary's trade-offs.

Prerequisites

Intuition

BPE training builds the vocabulary from the bottom up:

  1. Start with all the individual characters (or bytes).
  2. Count all the pairs of adjacent symbols in the corpus.
  3. Merge the most common pair into a new symbol.
  4. Repeat until the vocabulary is the size you want.

The result: common words become a single token, unusual ones are built from parts. Nothing is unknown — in the worst case the word is spelled out character by character.

Byte-level BPE (GPT-2 onwards) starts from bytes instead of characters. Then any text at all can be represented, including emoji and every language, with no <unk>.

Code

from collections import Counter

def bpe_train(word_frequency, n_merges):
    vocab = {tuple(w) + ("</w>",): f for w, f in word_frequency.items()}
    merges = []
    for _ in range(n_merges):
        pairs = Counter()
        for symbols, f in vocab.items():
            for a, b in zip(symbols, symbols[1:]):
                pairs[(a, b)] += f
        if not pairs:
            break
        best = pairs.most_common(1)[0][0]
        merges.append(best)
        new = {}
        for symbols, f in vocab.items():
            out, i = [], 0
            while i < len(symbols):
                if i + 1 < len(symbols) and (symbols[i], symbols[i+1]) == best:
                    out.append(symbols[i] + symbols[i+1]); i += 2
                else:
                    out.append(symbols[i]); i += 1
            new[tuple(out)] = f
        vocab = new
    return merges

corpus = {"låg": 5, "lågt": 3, "lägre": 2, "låga": 4}   # Swedish for low/lower
print(bpe_train(corpus, 4))
# [('l', 'å'), ('lå', 'g'), something like ('lägre', '</w>') depending on the frequencies]

The trade-offs with the vocabulary size:

Small (8 k)Large (200 k)
a small embedding matrixa large matrix (200k × d parameters)
long sequences → expensive attentionshort sequences → cheaper
every token is trained a lotrare tokens are barely trained

Swedish often gets 1.5–2× more tokens per word than English in models trained mostly on English — which means directly higher costs and a less effective context window for Swedish texts.

Mastery means

  • Implements BPE training and encoding
  • Explains the vocabulary's trade-offs
  • Sees how tokenisation affects cost and quality

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences