EUniversityTransformer architecture· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Attention variants: cross, causal, sparse
Be able to tell self-, cross- and causal attention apart and explain sparse and local variants.
Prerequisites
- EMulti-head attention in detailrequired
Intuition
By who supplies Q, K and V:
- Self-attention: all three from the same sequence.
- Cross-attention: Q from one sequence, K and V from another (the decoder ← the encoder).
By what may be seen:
- Full (bidirectional): all the positions see all of them. Encoders.
- Causal: position t sees only ≤ t. Decoders.
- Prefix: full within a prompt part, causal after it.
By how much is computed — the motive is that full attention costs O(T²):
- Local/sliding window: see only the w nearest (Mistral uses 4 096).
- Sparse: fixed patterns, for instance every other layer local and every other global (Longformer, BigBird).
- Linear: approximate the softmax so that the cost becomes O(T) (Performer).
- MQA/GQA: fewer K/V heads than Q heads — the same computation but a much smaller KV cache.
Code
import torch
T = 8
causal = torch.tril(torch.ones(T, T, dtype=torch.bool))
print(causal.int()[:4])
# [[1 0 0 0 0 0 0 0]
# [1 1 0 0 0 0 0 0]
# [1 1 1 0 0 0 0 0]
# [1 1 1 1 0 0 0 0]]
w = 3 # a sliding window, causal
i = torch.arange(T)[:, None]; j = torch.arange(T)[None, :]
local = (j <= i) & (j > i - w)
print(local.int()[5]) # [0 0 0 1 1 1 0 0] ← it sees only three back
# the cost: full against a window
for Tn in (1_000, 32_000):
print(Tn, f"full {Tn**2:,}", f"window(4096) {min(Tn, 4096) * Tn:,}")
# 1000 full 1,000,000 window 1,000,000
# 32000 full 1,024,000,000 window 131,072,000 ← 8× cheaper
The trade-off: a sliding window can only reach further back by stacking layers (after L layers the reach is L·w). That is often enough, but a model with only local attention can miss a dependency on something 20 000 tokens earlier. That is why most architectures mix local and global layers.
Mastery means
- Tells self-, cross- and causal attention apart
- Explains sparse and local variants and why they are needed
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Longformer: The Long-Document Transformer — arXiv (open access; licence per article)
- arXiv — GQA: Training Generalized Multi-Query Transformer Models — arXiv (open access; licence per article)