Skip to content
AI-grafen
GFrontier LabInterpretability· about 120 min· fast-moving, sources checked often· verified 2026-09-20· EN

The logit lens and the residual stream

Be able to project intermediate layers to the vocabulary and interpret the residual stream.

Prerequisites

Intuition

In a transformer every layer adds to the same vector — the residual stream. The output layer (the unembedding) translates the final vector into probabilities over the vocabulary.

The logit lens: apply the same unembedding to the residual stream after every layer, not just the last. The result is a sequence of «what would the model have answered if it had stopped here?».

The typical pattern: the first layers give nonsense or high-frequency words, somewhere in the middle the right answer appears, and the last layers sharpen the probability. Where the answer «arises» is often informative — especially when comparing prompts the model manages with ones it does not.

Code

from transformer_lens import HookedTransformer
import torch

m = HookedTransformer.from_pretrained("gpt2-small")
tokens = m.to_tokens("The capital of France is")
logits, cache = m.run_with_cache(tokens)

for layer in range(m.cfg.n_layers):
    resid = cache[f"blocks.{layer}.hook_resid_post"][0, -1]       # the last position
    resid = m.ln_final(resid[None, None])                          # the same normalisation as the output
    top = m.unembed(resid)[0, 0].softmax(-1).topk(3)
    words = [m.to_string(i) for i in top.indices]
    print(f"layer {layer:2d}: " + ", ".join(f"{w!r} {p:.2f}" for w, p in zip(words, top.values)))
# layer  0: ' the' 0.03, ',' 0.02, ' a' 0.02      ← high-frequency words
# layer  6: ' Paris' 0.08, ' France' 0.05, ...    ← the answer starts to emerge
# layer 11: ' Paris' 0.61, ' Lyon' 0.04, ...      ← sharpened

Two important reservations:

  1. It does not work equally well on every model. The logit lens presupposes that the intermediate layers' representations lie in roughly the same «space» as the output. In many models that holds badly, and the result is noise.
  2. The tuned lens (Belrose et al. 2023) solves that by learning an affine transformation per layer before the unembedding. It gives markedly more reliable and comparable results — use it when the logit lens looks messy.

Mastery means

  • Projects intermediate layers to the vocabulary
  • Interprets the residual stream layer by layer
  • Knows the method's limitations and the tuned lens

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences