Saliency and attribution
Be able to compute gradient-based attribution and know its shortcomings.
Prerequisites
Intuition
Attribution answers: which parts of the input affected the output most?
| The method | The idea | The problem |
|---|---|---|
| Vanilla saliency | ∂output/∂input | |
| Input × Gradient | the gradient weighted by the input | still noisy |
| Integrated Gradients | integrate the gradient along a path from a baseline | it requires choosing a baseline |
| Grad-CAM | weight the last convolutional layer's activation maps | CNNs only, a coarse resolution |
| SHAP | Shapley values, axiomatically grounded | expensive, approximated in practice |
Integrated Gradients is often the first choice because it satisfies two reasonable axioms: completeness (the attributions sum to the difference from the baseline) and sensitivity.
Formal
where is a baseline (a black image, the zero vector, the average input). The integral is approximated with 20–300 steps.
The choice of baseline is not neutral: a black image gives zero attribution to black pixels — which are thereby made invisible. For text a sequence of padding tokens is often used, which has the same problem.
Sanity checks (Adebayo et al. 2018) — always run these:
- Model randomisation: randomise the model's weights layer by layer. A valid attribution method should give completely different maps. Several popular methods do not — they essentially produce edge detection whatever the model.
- Label randomisation: retrain the model on random labels. The attributions should change.
A method that passes both is at least connected to the model. One that does not shows something about the image, not about the decision.
The practical consequence: use attribution to generate hypotheses («the model seems to be looking at the background»), and then verify with an intervention («mask the background and measure»). Never as a final explanation — and never as legal evidence.
Code
import torch
def integrated_gradients(model, x, target, baseline=None, steps=64):
baseline = torch.zeros_like(x) if baseline is None else baseline
alphas = torch.linspace(0, 1, steps).view(-1, *([1] * x.dim()))
path = baseline + alphas * (x - baseline) # (steps, ...)
path.requires_grad_(True)
out = model(path)[:, target].sum()
grad = torch.autograd.grad(out, path)[0]
mean_grad = grad.mean(dim=0)
ig = (x - baseline) * mean_grad
# the completeness check: the sum should ≈ f(x) − f(baseline)
with torch.no_grad():
diff = (model(x[None])[0, target] - model(baseline[None])[0, target]).item()
return ig, {"ig_sum": ig.sum().item(), "f_diff": diff}
ig, check = integrated_gradients(model, image, target=cls)
print(check) # {'ig_sum': 3.81, 'f_diff': 3.94} ← close = the approximation is good enough
The completeness check is free and reveals straight away whether the number of steps is too small.
Mastery means
- Computes gradient-based attribution
- Knows the methods' shortcomings
- Uses sanity checks
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Axiomatic Attribution for Deep Networks (Integrated Gradients) — arXiv (open access; licence per article)
- arXiv — Sanity Checks for Saliency Maps — arXiv (open access; licence per article)