Byte-pair encoding
Be able to implement BPE training and tokenisation and explain the vocabulary's trade-offs.
Prerequisites
- DTokenisationrequired
Intuition
BPE training builds the vocabulary from the bottom up:
- Start with all the individual characters (or bytes).
- Count all the pairs of adjacent symbols in the corpus.
- Merge the most common pair into a new symbol.
- Repeat until the vocabulary is the size you want.
The result: common words become a single token, unusual ones are built from parts. Nothing is unknown — in the worst case the word is spelled out character by character.
Byte-level BPE (GPT-2 onwards) starts from bytes instead of characters. Then any text at all can be represented, including emoji and every language, with no <unk>.
Code
from collections import Counter
def bpe_train(word_frequency, n_merges):
vocab = {tuple(w) + ("</w>",): f for w, f in word_frequency.items()}
merges = []
for _ in range(n_merges):
pairs = Counter()
for symbols, f in vocab.items():
for a, b in zip(symbols, symbols[1:]):
pairs[(a, b)] += f
if not pairs:
break
best = pairs.most_common(1)[0][0]
merges.append(best)
new = {}
for symbols, f in vocab.items():
out, i = [], 0
while i < len(symbols):
if i + 1 < len(symbols) and (symbols[i], symbols[i+1]) == best:
out.append(symbols[i] + symbols[i+1]); i += 2
else:
out.append(symbols[i]); i += 1
new[tuple(out)] = f
vocab = new
return merges
corpus = {"låg": 5, "lågt": 3, "lägre": 2, "låga": 4} # Swedish for low/lower
print(bpe_train(corpus, 4))
# [('l', 'å'), ('lå', 'g'), something like ('lägre', '</w>') depending on the frequencies]
The trade-offs with the vocabulary size:
| Small (8 k) | Large (200 k) |
|---|---|
| a small embedding matrix | a large matrix (200k × d parameters) |
| long sequences → expensive attention | short sequences → cheaper |
| every token is trained a lot | rare tokens are barely trained |
Swedish often gets 1.5–2× more tokens per word than English in models trained mostly on English — which means directly higher costs and a less effective context window for Swedish texts.
Mastery means
- Implements BPE training and encoding
- Explains the vocabulary's trade-offs
- Sees how tokenisation affects cost and quality
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Neural Machine Translation of Rare Words with Subword Units — arXiv (open access; licence per article)
- Hugging Face — dokumentation (Apache-2.0) — Apache-2.0