Skip to content
AI-grafen
EUniversityTransformer architecture· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Attention variants: cross, causal, sparse

Be able to tell self-, cross- and causal attention apart and explain sparse and local variants.

Prerequisites

Intuition

By who supplies Q, K and V:

  • Self-attention: all three from the same sequence.
  • Cross-attention: Q from one sequence, K and V from another (the decoder ← the encoder).

By what may be seen:

  • Full (bidirectional): all the positions see all of them. Encoders.
  • Causal: position t sees only ≤ t. Decoders.
  • Prefix: full within a prompt part, causal after it.

By how much is computed — the motive is that full attention costs O(T²):

  • Local/sliding window: see only the w nearest (Mistral uses 4 096).
  • Sparse: fixed patterns, for instance every other layer local and every other global (Longformer, BigBird).
  • Linear: approximate the softmax so that the cost becomes O(T) (Performer).
  • MQA/GQA: fewer K/V heads than Q heads — the same computation but a much smaller KV cache.

Code

import torch

T = 8
causal = torch.tril(torch.ones(T, T, dtype=torch.bool))
print(causal.int()[:4])
# [[1 0 0 0 0 0 0 0]
#  [1 1 0 0 0 0 0 0]
#  [1 1 1 0 0 0 0 0]
#  [1 1 1 1 0 0 0 0]]

w = 3                                      # a sliding window, causal
i = torch.arange(T)[:, None]; j = torch.arange(T)[None, :]
local = (j <= i) & (j > i - w)
print(local.int()[5])                      # [0 0 0 1 1 1 0 0]  ← it sees only three back

# the cost: full against a window
for Tn in (1_000, 32_000):
    print(Tn, f"full {Tn**2:,}", f"window(4096) {min(Tn, 4096) * Tn:,}")
# 1000   full 1,000,000       window 1,000,000
# 32000  full 1,024,000,000   window 131,072,000    ← 8× cheaper

The trade-off: a sliding window can only reach further back by stacking layers (after L layers the reach is L·w). That is often enough, but a model with only local attention can miss a dependency on something 20 000 tokens earlier. That is why most architectures mix local and global layers.

Mastery means

  • Tells self-, cross- and causal attention apart
  • Explains sparse and local variants and why they are needed

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences