The logit lens and the residual stream
Be able to project intermediate layers to the vocabulary and interpret the residual stream.
Prerequisites
Intuition
In a transformer every layer adds to the same vector — the residual stream. The output layer (the unembedding) translates the final vector into probabilities over the vocabulary.
The logit lens: apply the same unembedding to the residual stream after every layer, not just the last. The result is a sequence of «what would the model have answered if it had stopped here?».
The typical pattern: the first layers give nonsense or high-frequency words, somewhere in the middle the right answer appears, and the last layers sharpen the probability. Where the answer «arises» is often informative — especially when comparing prompts the model manages with ones it does not.
Code
from transformer_lens import HookedTransformer
import torch
m = HookedTransformer.from_pretrained("gpt2-small")
tokens = m.to_tokens("The capital of France is")
logits, cache = m.run_with_cache(tokens)
for layer in range(m.cfg.n_layers):
resid = cache[f"blocks.{layer}.hook_resid_post"][0, -1] # the last position
resid = m.ln_final(resid[None, None]) # the same normalisation as the output
top = m.unembed(resid)[0, 0].softmax(-1).topk(3)
words = [m.to_string(i) for i in top.indices]
print(f"layer {layer:2d}: " + ", ".join(f"{w!r} {p:.2f}" for w, p in zip(words, top.values)))
# layer 0: ' the' 0.03, ',' 0.02, ' a' 0.02 ← high-frequency words
# layer 6: ' Paris' 0.08, ' France' 0.05, ... ← the answer starts to emerge
# layer 11: ' Paris' 0.61, ' Lyon' 0.04, ... ← sharpened
Two important reservations:
- It does not work equally well on every model. The logit lens presupposes that the intermediate layers' representations lie in roughly the same «space» as the output. In many models that holds badly, and the result is noise.
- The tuned lens (Belrose et al. 2023) solves that by learning an affine transformation per layer before the unembedding. It gives markedly more reliable and comparable results — use it when the logit lens looks messy.
Mastery means
- Projects intermediate layers to the vocabulary
- Interprets the residual stream layer by layer
- Knows the method's limitations and the tuned lens
Sign in to do the exercises and build your mastery up.
Sources
- nostalgebraist — interpreting GPT: the logit lens — free to read
- arXiv — Eliciting Latent Predictions from Transformers with the Tuned Lens — arXiv (open access; licence per article)