Skip to content
AI-grafen
DAI developerTransformer architecture· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Transformers — the architecture

Be able to draw and explain a transformer block (attention, residual, layernorm, MLP), explain why positional information is needed, and describe how a decoder-only model generates text token by token.

Prerequisites

Intuition

A transformer block does two things with every token: it lets the token gather information from other tokens (attention) and then process it on its own (a small neural network, the MLP). Around both there is a residual connection (add the input back in, so nothing is lost) and layernorm (keep the numbers on a sensible scale).

Stack 12, 32 or 96 blocks like that and you have the GPT, Llama or Qwen family.

Formal

For token representations X (n×d):

X ← X + MultiHeadAttention(LayerNorm(X)) X ← X + MLP(LayerNorm(X))

Multi-head: h parallel attention heads of d/h dimensions each, the results concatenated — different heads can learn different kinds of relationship.

Attention is permutation-invariant — it does not know the order. That is why positional encoding is added to the embeddings (sinusoids in the original, learnt or rotary (RoPE) in modern models).

Decoder-only (GPT style): a causal mask so that token t only sees 1…t. The last layer gives logits over the vocabulary → softmax → the probability of the next token. Generation: draw a token, append it, run again. One token at a time, which is why generation is slow and the KV cache (saving K and V for the earlier tokens) is decisive for speed.

Code

import torch.nn as nn

class Block(nn.Module):
    def __init__(self, d, heads):
        super().__init__()
        self.ln1, self.ln2 = nn.LayerNorm(d), nn.LayerNorm(d)
        self.attn = nn.MultiheadAttention(d, heads, batch_first=True)
        self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))

    def forward(self, x, mask):
        h = self.ln1(x)
        x = x + self.attn(h, h, h, attn_mask=mask)[0]
        return x + self.mlp(self.ln2(x))

mask is an n×n upper triangular matrix with −inf above the diagonal: the causal mask.

Mastery means

  • Describes the components of a transformer block and their order
  • Explains autoregressive generation and the causal mask

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences