Skip to content
AI-grafen
GFrontier LabInterpretability· about 120 min· fast-moving, sources checked often· verified 2026-09-20· EN

Circuits, ablation and activation patching

Be able to find a circuit for a simple behaviour and verify it with interventions.

Prerequisites

Intuition

Activation patching: run the model on a «clean» prompt and a «corrupt» one (the names swapped, say). Take the activation from a component (a head, a layer, a position) in the clean run and paste it into the corrupt one. If the output now becomes «clean» — the component carries the information. Do it for every component → a map of what matters.

Ablation: zero (or replace with the mean) a component and measure how much of the behaviour disappears. Mean ablation is gentler than zero ablation (zero is outside the distribution).

A circuit = the smallest set of components that suffices for the behaviour: verify it by running the model with only the circuit (the rest mean-ablated) and measuring how much of the effect remains (80 % of the logit difference, say). Report the completeness (how much the circuit explains) and the minimality (that every part is needed).

The measure: the logit difference between the right and the wrong answer is more stable than the probability.

Code

from transformer_lens import HookedTransformer
import torch

m = HookedTransformer.from_pretrained("gpt2-small")
clean = m.to_tokens("When Mary and John went to the store, John gave a drink to")
corrupt = m.to_tokens("When Mary and John went to the store, Mary gave a drink to")
MARY, JOHN = m.to_single_token(" Mary"), m.to_single_token(" John")

def logit_diff(logits):
    return (logits[0, -1, MARY] - logits[0, -1, JOHN]).item()

_, cache_clean = m.run_with_cache(clean)
base_clean, base_corr = logit_diff(m(clean)), logit_diff(m(corrupt))

def patch_head(layer, head):
    def hook(z, hook):                     # z: (batch, pos, head, d_head)
        z[:, :, head, :] = cache_clean[hook.name][:, :, head, :]
        return z
    logits = m.run_with_hooks(corrupt, fwd_hooks=[(f"blocks.{layer}.attn.hook_z", hook)])
    return (logit_diff(logits) - base_corr) / (base_clean - base_corr)   # 1 = fully restored

res = torch.tensor([[patch_head(l, h) for h in range(12)] for l in range(12)])
print(res.topk(5))                          # the heads that restore the most — circuit candidates

# verify: mean-ablate EVERYTHING but the candidates, measure the remaining logit diff

Research

The methodology was formalised in Wang et al. (2022) (IOI), Meng et al. (2022) (causal tracing/ROME for factual memory) and Conmy et al. (2023) (ACDC — automatic circuit discovery). Important warnings: patching at one position can give misleading results under redundancy (backup heads take over when one is ablated — Wang 2022 found exactly this); a distribution shift at zero ablation; and that attribution methods (gradient-based) can differ from intervention-based ones. Causal scrubbing (Chan et al. 2022) is a stricter hypothesis test: replace activations with ones from inputs the hypothesis says should be equivalent. The field is moving towards sparse autoencoder features as the units instead of heads and neurons, and towards circuits in production models (Anthropic 2024–2025, attribution graphs).

Mastery means

  • Performs activation patching and interprets the effect
  • Identifies a circuit for a simple behaviour and verifies it with ablation
  • Reports with controls and a measure of how much of the behaviour the circuit explains

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences