Mechanistic interpretability — the basics
Be able to inspect activations and attention heads, find circuits for simple behaviours and verify them with interventions.
Prerequisites
- DTransformers — the architecturerequired
- EMulti-head attention in detailrequired
- EScientific method in AIrequired
Intuition
Mechanistic interpretability tries to understand what the network computes — not just that it works. The tools:
- The residual stream: in a transformer every layer reads from and writes to the same vector per token. It can be read off after every layer.
- The logit lens: project the residual stream after layer ℓ through the unembedding → which token does the model «believe» in halfway through? The answer can often be seen emerging layer by layer.
- Attention patterns: some heads do clear things: «the previous token», «copy from an earlier occurrence» (induction heads — the basis of in-context learning).
- Probes: train a linear classifier on the activations to see whether a property (for instance «this is a verb») is linearly readable.
But: observation ≠ mechanism. That a head looks at the subject does not prove that it is used. Hence: intervention — change the activation and see whether the output changes as the hypothesis predicts. That is the next node.
Code
# With TransformerLens (Nanda) on GPT-2 small
from transformer_lens import HookedTransformer
import torch
m = HookedTransformer.from_pretrained("gpt2-small")
text = "When Mary and John went to the store, John gave a drink to"
toks = m.to_tokens(text)
logits, cache = m.run_with_cache(toks)
print(m.to_string(logits[0, -1].argmax())) # ' Mary'
# the logit lens: what does the model believe after each layer?
for l in range(m.cfg.n_layers):
resid = cache[f"blocks.{l}.hook_resid_post"][0, -1]
top = m.unembed(m.ln_final(resid[None, None]))[0, 0].argmax()
print(l, repr(m.to_string(top)))
# the attention pattern for head (9, 9) — a known «name mover» in the IOI circuit
attn = cache["blocks.9.attn.hook_pattern"][0, 9] # (T, T)
print(attn[-1].topk(3)) # does the last token look at ' Mary'?
# a linear probe: is «the token is a name» readable in layer 6?
# X = cache["blocks.6.hook_resid_post"][0]; y = labels; LogisticRegression().fit(X, y)
Research
The field builds on Elhage et al. (2021), A Mathematical Framework for Transformer Circuits: attention heads as read/write operations on the residual stream, the QK circuit (where you look) and the OV circuit (what is copied). Olsson et al. (2022) showed induction heads and their connection to in-context learning. Wang et al. (2022) mapped the IOI circuit in GPT-2 small (26 heads in seven roles) with activation patching. Superposition (Elhage 2022) explains why individual neurons are polysemantic, and sparse autoencoders (Bricken 2023, Templeton 2024) are used to find interpretable «features» in large models. Open questions: scaling the methods, how much of the behaviour is explained by the circuits found, and whether interpretations are faithful (causally correct) or merely plausible.
Mastery means
- Inspects activations, the residual stream and attention heads in a small transformer
- Uses the logit lens and attention patterns for hypotheses
- Verifies hypotheses with interventions, not just observation
Sign in to do the exercises and build your mastery up.
Sources
- Elhage m.fl. — A Mathematical Framework for Transformer Circuits (Anthropic) — free to read
- arXiv — In-context Learning and Induction Heads — arXiv (open access; licence per article)
- TransformerLens (MIT) — MIT