Skip to content
AI-grafen
FAI engineeringInterpretability· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Activations and linear probes

Be able to train a linear probe on activations and interpret what a layer represents.

Prerequisites

Intuition

A linear probe is a simple classifier (often logistic regression) trained on a layer's activations to predict a property: is the word a verb? is the statement true? which language is the text in?

If the probe succeeds, the property is linearly readable in that layer. That is a useful signal about what the representation carries.

But: that the information is there does not mean the model uses it. A probe measures readability, not causality. For causality you need an intervention (activation patching).

Formal

The probe's fundamental problem: a sufficiently expressive probe can learn the property itself from the activations, even if the model does not represent it. Hence:

  1. Use a linear probe — limited capacity, it measures whether the information is easily accessible.
  2. Run a control task (Hewitt & Liang 2019): train the same probe on randomly reassigned labels. High «selectivity» = a large difference between the real and the control task → the probe is reading the model, not memorising.
  3. Compare against baselines: the same probe on random weights (an untrained model) and on the embedding layer. If layer 12 does not beat the embedding layer, the result says nothing about what the model has learnt.
  4. Report a layer profile, not a single layer — where the information arises and where it disappears is the result itself.

Amnestic probing goes a step further: remove the direction the probe found (with INLP, for instance) and measure whether the model's behaviour changes. If it does not, the information was not being used.

Code

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

def probe_per_layer(activations: dict[int, np.ndarray], y: np.ndarray, seed=0):
    """activations: {layer: (n, d)}. Returns the accuracy and the selectivity per layer."""
    rng = np.random.default_rng(seed)
    y_control = rng.permutation(y)                       # the control task
    rows = []
    for layer, X in sorted(activations.items()):
        clf = LogisticRegression(max_iter=2000, C=1.0)
        acc = cross_val_score(clf, X, y, cv=5).mean()
        ctrl = cross_val_score(clf, X, y_control, cv=5).mean()
        rows.append({"layer": layer, "acc": round(acc, 3),
                     "control": round(ctrl, 3), "selectivity": round(acc - ctrl, 3)})
    return rows

for r in probe_per_layer(act, y):
    print(r)
# {'layer': 0,  'acc': 0.62, 'control': 0.50, 'selectivity': 0.12}
# {'layer': 6,  'acc': 0.88, 'control': 0.51, 'selectivity': 0.37}   ← the information arises here
# {'layer': 11, 'acc': 0.91, 'control': 0.50, 'selectivity': 0.41}

The layer profile is the result: it shows where in the network the property becomes readable — and the control column shows that the probe is not just memorising.

Mastery means

  • Trains a linear probe on activations
  • Interprets the result with the right reservations
  • Uses controls against false conclusions

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences