Skip to content
AI-grafen
DAI developerInterpretability· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Looking inside the network

Be able to visualise what individual neurons respond to in a small network.

Prerequisites

Intuition

A trained network is full of numbers. The question «what does this neuron do?» can actually be investigated, and there are three standard methods:

MethodHowGives
Maximally activating examplesrun the whole dataset, keep the inputs that gave the highest activationsimple, honest, always a good first step
Feature visualisationoptimise a synthetic image that maximises the neuronbeautiful pictures, but they can show things that do not exist in reality
Ablationzero the neuron out and measure what changesthe only one that shows cause

Always start with the first. It requires no optimisation, cannot make anything up, and answers the question you actually asked.

But be careful with the interpretation. That a neuron is activated by dog pictures does not mean it «is» a dog detector. It may be responding to fur, to a grass background, or to something that happens to co-vary with dogs in your particular dataset.

Formal

Polysemantic neurons are the big obstacle. One and the same neuron can respond to apparently unrelated things — cat faces, car fronts and the letter A.

The explanation is superposition: the network needs to represent more concepts than it has neurons, and therefore packs several into each. Concepts that rarely occur at the same time can share a neuron without disturbing each other much.

That means individual neurons are often the wrong unit to study. Directions in activation space — combinations of neurons — are more often monosemantic. That is the whole idea behind sparse autoencoders (SAEs), which learn a larger but sparser basis where each direction corresponds to a concept.

Three traps:

TrapWhy
Only looking at the top-9 imagesthe strongest activations are not representative — look at the middle of the distribution too
Drawing conclusions without an interventioncorrelation; ablate the neuron and see whether anything actually changes
Taking feature visualisation literallythe optimised image lies outside the data distribution

A complete claim about a neuron has three parts:

  1. What activates it — maximally activating examples, from across the distribution.
  2. What does not — negative controls that would have activated it if your hypothesis were wrong in a particular way.
  3. What happens without it — ablation, with the effect on the output measured.

Without part 3 you have described a correlation, not a function.

Code

import torch, torch.nn as nn

def maximally_activating(model, layer, neuron, dataloader, k=9):
    """Find the k inputs that activate a neuron the most — and a few from the middle."""
    stored = {}
    h = layer.register_forward_hook(lambda m, i, o: stored.update(a=o.detach()))
    scores, images = [], []
    with torch.no_grad():
        for x, _ in dataloader:
            model(x)
            a = stored["a"]
            v = a[:, neuron].flatten(1).max(dim=1).values if a.dim() == 4 else a[:, neuron]
            scores += v.tolist(); images += list(x)
    h.remove()
    order = sorted(range(len(scores)), key=lambda i: -scores[i])
    middle = order[len(order) // 2:len(order) // 2 + k]
    return ([images[i] for i in order[:k]], [images[i] for i in middle],
            [round(scores[i], 3) for i in order[:k]])

def ablate(model, layer, neuron, dataloader):
    """Zero the neuron out and measure how much the accuracy changes."""
    def zero(m, i, o):
        o = o.clone(); o[:, neuron] = 0; return o

    def accuracy():
        right = n = 0
        with torch.no_grad():
            for x, y in dataloader:
                right += int((model(x).argmax(1) == y).sum()); n += len(y)
        return right / n

    before = accuracy()
    h = layer.register_forward_hook(zero)
    after = accuracy()
    h.remove()
    return {"before": round(before, 4), "after": round(after, 4),
            "effect": round(before - after, 4)}

# print(ablate(model, model[3], neuron=7, dataloader=vl))
# {'before': 0.9881, 'after': 0.9878, 'effect': 0.0003}
#  ← the neuron is redundant: the network manages without it

The outcome in the comment is the most common and the most instructive: ablating one neuron changes almost nothing, since information is spread over many. That is in itself an important result — and a reason to be careful with stories about what individual neurons «do».

Mastery means

  • Finds which inputs activate a neuron the most
  • Interprets the result carefully
  • Knows about polysemantic neurons

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences