Transformers — the architecture
Be able to draw and explain a transformer block (attention, residual, layernorm, MLP), explain why positional information is needed, and describe how a decoder-only model generates text token by token.
Prerequisites
- DAttentionrequired
- DTrain a neural network in PyTorchrequired
Intuition
A transformer block does two things with every token: it lets the token gather information from other tokens (attention) and then process it on its own (a small neural network, the MLP). Around both there is a residual connection (add the input back in, so nothing is lost) and layernorm (keep the numbers on a sensible scale).
Stack 12, 32 or 96 blocks like that and you have the GPT, Llama or Qwen family.
Formal
For token representations X (n×d):
X ← X + MultiHeadAttention(LayerNorm(X)) X ← X + MLP(LayerNorm(X))
Multi-head: h parallel attention heads of d/h dimensions each, the results concatenated — different heads can learn different kinds of relationship.
Attention is permutation-invariant — it does not know the order. That is why positional encoding is added to the embeddings (sinusoids in the original, learnt or rotary (RoPE) in modern models).
Decoder-only (GPT style): a causal mask so that token t only sees 1…t. The last layer gives logits over the vocabulary → softmax → the probability of the next token. Generation: draw a token, append it, run again. One token at a time, which is why generation is slow and the KV cache (saving K and V for the earlier tokens) is decisive for speed.
Code
import torch.nn as nn
class Block(nn.Module):
def __init__(self, d, heads):
super().__init__()
self.ln1, self.ln2 = nn.LayerNorm(d), nn.LayerNorm(d)
self.attn = nn.MultiheadAttention(d, heads, batch_first=True)
self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
def forward(self, x, mask):
h = self.ln1(x)
x = x + self.attn(h, h, h, attn_mask=mask)[0]
return x + self.mlp(self.ln2(x))
mask is an n×n upper triangular matrix with −inf above the diagonal: the causal mask.
Mastery means
- Describes the components of a transformer block and their order
- Explains autoregressive generation and the causal mask
Sign in to do the exercises and build your mastery up.
Sources
Leads to
- EBERT and masked language modelling
- EEncoder–decoder transformers
- ELayerNorm, RMSNorm and pre-/post-norm
- EThe MLP block: GELU, SwiGLU
- EMulti-head attention in detail
- EPositional encoding and RoPE
- ELanguage models — training and generation
- EVision Transformer (ViT)
- FInference optimisation
- FReading and analysing research papers
- FSpeech recognition (ASR)
- GMechanistic interpretability — the basics
- GState space models and Mamba
Part of the goals (25)
- Build a transformer from scratch
- Understand how generative AI works
- Build a voice interface
- Run models more cheaply: quantisation
- Frontier Lab — an independent research project
- Language models in practice
- Evals in practice
- Reproduce a paper
- Interpreting a language model
- Multimodal systems
- Fine-tune and run your own models
- Build a RAG system you can trust
- AI safety in practice
- Fine-tune a model with LoRA
- Responsible AI in practice
- Build an agent you can trust
- Build an NLP system end to end
- AI in production
- Generative models in depth
- An AI service in operation
- Build a memory system for an agent
- Build an AI service that survives production
- Deep reinforcement learning
- Statistics for experiments
- AI, ethics and society