Looking inside the network
Be able to visualise what individual neurons respond to in a small network.
Prerequisites
Intuition
A trained network is full of numbers. The question «what does this neuron do?» can actually be investigated, and there are three standard methods:
| Method | How | Gives |
|---|---|---|
| Maximally activating examples | run the whole dataset, keep the inputs that gave the highest activation | simple, honest, always a good first step |
| Feature visualisation | optimise a synthetic image that maximises the neuron | beautiful pictures, but they can show things that do not exist in reality |
| Ablation | zero the neuron out and measure what changes | the only one that shows cause |
Always start with the first. It requires no optimisation, cannot make anything up, and answers the question you actually asked.
But be careful with the interpretation. That a neuron is activated by dog pictures does not mean it «is» a dog detector. It may be responding to fur, to a grass background, or to something that happens to co-vary with dogs in your particular dataset.
Formal
Polysemantic neurons are the big obstacle. One and the same neuron can respond to apparently unrelated things — cat faces, car fronts and the letter A.
The explanation is superposition: the network needs to represent more concepts than it has neurons, and therefore packs several into each. Concepts that rarely occur at the same time can share a neuron without disturbing each other much.
That means individual neurons are often the wrong unit to study. Directions in activation space — combinations of neurons — are more often monosemantic. That is the whole idea behind sparse autoencoders (SAEs), which learn a larger but sparser basis where each direction corresponds to a concept.
Three traps:
| Trap | Why |
|---|---|
| Only looking at the top-9 images | the strongest activations are not representative — look at the middle of the distribution too |
| Drawing conclusions without an intervention | correlation; ablate the neuron and see whether anything actually changes |
| Taking feature visualisation literally | the optimised image lies outside the data distribution |
A complete claim about a neuron has three parts:
- What activates it — maximally activating examples, from across the distribution.
- What does not — negative controls that would have activated it if your hypothesis were wrong in a particular way.
- What happens without it — ablation, with the effect on the output measured.
Without part 3 you have described a correlation, not a function.
Code
import torch, torch.nn as nn
def maximally_activating(model, layer, neuron, dataloader, k=9):
"""Find the k inputs that activate a neuron the most — and a few from the middle."""
stored = {}
h = layer.register_forward_hook(lambda m, i, o: stored.update(a=o.detach()))
scores, images = [], []
with torch.no_grad():
for x, _ in dataloader:
model(x)
a = stored["a"]
v = a[:, neuron].flatten(1).max(dim=1).values if a.dim() == 4 else a[:, neuron]
scores += v.tolist(); images += list(x)
h.remove()
order = sorted(range(len(scores)), key=lambda i: -scores[i])
middle = order[len(order) // 2:len(order) // 2 + k]
return ([images[i] for i in order[:k]], [images[i] for i in middle],
[round(scores[i], 3) for i in order[:k]])
def ablate(model, layer, neuron, dataloader):
"""Zero the neuron out and measure how much the accuracy changes."""
def zero(m, i, o):
o = o.clone(); o[:, neuron] = 0; return o
def accuracy():
right = n = 0
with torch.no_grad():
for x, y in dataloader:
right += int((model(x).argmax(1) == y).sum()); n += len(y)
return right / n
before = accuracy()
h = layer.register_forward_hook(zero)
after = accuracy()
h.remove()
return {"before": round(before, 4), "after": round(after, 4),
"effect": round(before - after, 4)}
# print(ablate(model, model[3], neuron=7, dataloader=vl))
# {'before': 0.9881, 'after': 0.9878, 'effect': 0.0003}
# ← the neuron is redundant: the network manages without it
The outcome in the comment is the most common and the most instructive: ablating one neuron changes almost nothing, since information is spread over many. That is in itself an important result — and a reason to be careful with stories about what individual neurons «do».
Mastery means
- Finds which inputs activate a neuron the most
- Interprets the result carefully
- Knows about polysemantic neurons
Sign in to do the exercises and build your mastery up.
Sources
- Distill — Feature Visualization (CC BY 4.0) — CC BY 4.0
- arXiv — Toy Models of Superposition — arXiv (open access; licence per article)
- PyTorch — tutorials (BSD-3) — BSD-3-Clause