The transformer — an overview without formulas
Be able to describe the flow token → embedding → attention → next word, in words and in a picture.
Prerequisites
Intuition
What happens when you type a question to a language model? Five steps:
1. Tokenisation. The text is split into pieces: «The cat sleeps» → ["The", " cat", " sleeps"] → numbers [8241, 268, 23409].
2. Embedding. Every token becomes a vector — a point on a map with hundreds of dimensions, where words with similar meanings sit close together. Positional information is added, otherwise the model does not know the order.
3. Attention, many times over. Every token looks at all the earlier ones and fetches what it needs. «It» looks at «the cat». This is repeated layer after layer — 32 layers in a typical model — so that every token carries more and more context.
4. Output. The last token's vector is multiplied against the vocabulary → a score per possible next word → softmax → probabilities.
5. Choosing and repeating. One word is chosen (the most likely, or sampled according to the probabilities), appended, and the whole thing starts over.
Interactive
Follow an example: «The capital of Sweden is»
| Step | What happens |
|---|---|
| tokens | ["The", " capital", " of", " Sweden", " is"] |
| embeddings | five vectors, each of ~4 000 numbers |
| attention layers 1–8 | « is» starts fetching from « Sweden» and « capital» |
| layers 9–24 | the representation becomes more and more «which country, what kind of answer» |
| output | Stockholm 0.92 · a 0.02 · Gothenburg 0.01 … |
| choice | «Stockholm» is appended, everything runs again for the next word |
Two insights worth carrying with you:
- The model computes all the way through for every single word it writes. A 100-word answer = 100 passes.
- There is no database of facts. «Stockholm» comes out because the weights, trained on enormous amounts of text, make that particular word the most likely. Which is why it can be wrong — and why being wrong sounds just as confident as being right.
Mastery means
- Describes the flow token → embedding → attention → next word
- Explains what each step adds
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Attention Is All You Need — arXiv (open access; licence per article)
- Jay Alammar — The Illustrated Transformer — free to read