Activations and linear probes
Be able to train a linear probe on activations and interpret what a layer represents.
Prerequisites
Intuition
A linear probe is a simple classifier (often logistic regression) trained on a layer's activations to predict a property: is the word a verb? is the statement true? which language is the text in?
If the probe succeeds, the property is linearly readable in that layer. That is a useful signal about what the representation carries.
But: that the information is there does not mean the model uses it. A probe measures readability, not causality. For causality you need an intervention (activation patching).
Formal
The probe's fundamental problem: a sufficiently expressive probe can learn the property itself from the activations, even if the model does not represent it. Hence:
- Use a linear probe — limited capacity, it measures whether the information is easily accessible.
- Run a control task (Hewitt & Liang 2019): train the same probe on randomly reassigned labels. High «selectivity» = a large difference between the real and the control task → the probe is reading the model, not memorising.
- Compare against baselines: the same probe on random weights (an untrained model) and on the embedding layer. If layer 12 does not beat the embedding layer, the result says nothing about what the model has learnt.
- Report a layer profile, not a single layer — where the information arises and where it disappears is the result itself.
Amnestic probing goes a step further: remove the direction the probe found (with INLP, for instance) and measure whether the model's behaviour changes. If it does not, the information was not being used.
Code
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
def probe_per_layer(activations: dict[int, np.ndarray], y: np.ndarray, seed=0):
"""activations: {layer: (n, d)}. Returns the accuracy and the selectivity per layer."""
rng = np.random.default_rng(seed)
y_control = rng.permutation(y) # the control task
rows = []
for layer, X in sorted(activations.items()):
clf = LogisticRegression(max_iter=2000, C=1.0)
acc = cross_val_score(clf, X, y, cv=5).mean()
ctrl = cross_val_score(clf, X, y_control, cv=5).mean()
rows.append({"layer": layer, "acc": round(acc, 3),
"control": round(ctrl, 3), "selectivity": round(acc - ctrl, 3)})
return rows
for r in probe_per_layer(act, y):
print(r)
# {'layer': 0, 'acc': 0.62, 'control': 0.50, 'selectivity': 0.12}
# {'layer': 6, 'acc': 0.88, 'control': 0.51, 'selectivity': 0.37} ← the information arises here
# {'layer': 11, 'acc': 0.91, 'control': 0.50, 'selectivity': 0.41}
The layer profile is the result: it shows where in the network the property becomes readable — and the control column shows that the probe is not just memorising.
Mastery means
- Trains a linear probe on activations
- Interprets the result with the right reservations
- Uses controls against false conclusions
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Designing and Interpreting Probes with Control Tasks — arXiv (open access; licence per article)
- arXiv — Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals — arXiv (open access; licence per article)