Skip to content
AI-grafen
FAI engineeringInterpretability· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Saliency and attribution

Be able to compute gradient-based attribution and know its shortcomings.

Prerequisites

Intuition

Attribution answers: which parts of the input affected the output most?

The methodThe ideaThe problem
Vanilla saliency∂output/∂input
Input × Gradientthe gradient weighted by the inputstill noisy
Integrated Gradientsintegrate the gradient along a path from a baselineit requires choosing a baseline
Grad-CAMweight the last convolutional layer's activation mapsCNNs only, a coarse resolution
SHAPShapley values, axiomatically groundedexpensive, approximated in practice

Integrated Gradients is often the first choice because it satisfies two reasonable axioms: completeness (the attributions sum to the difference from the baseline) and sensitivity.

Formal

IGi(x)=(xi−xi′)∫α=01∂f(x′+α(x−x′))∂xi dα\text{IG}_i(x) = (x_i - x'_i)\int_{\alpha=0}^{1} \frac{\partial f(x' + \alpha(x-x'))}{\partial x_i}\,d\alpha

where x′x' is a baseline (a black image, the zero vector, the average input). The integral is approximated with 20–300 steps.

The choice of baseline is not neutral: a black image gives zero attribution to black pixels — which are thereby made invisible. For text a sequence of padding tokens is often used, which has the same problem.

Sanity checks (Adebayo et al. 2018) — always run these:

  1. Model randomisation: randomise the model's weights layer by layer. A valid attribution method should give completely different maps. Several popular methods do not — they essentially produce edge detection whatever the model.
  2. Label randomisation: retrain the model on random labels. The attributions should change.

A method that passes both is at least connected to the model. One that does not shows something about the image, not about the decision.

The practical consequence: use attribution to generate hypotheses («the model seems to be looking at the background»), and then verify with an intervention («mask the background and measure»). Never as a final explanation — and never as legal evidence.

Code

import torch

def integrated_gradients(model, x, target, baseline=None, steps=64):
    baseline = torch.zeros_like(x) if baseline is None else baseline
    alphas = torch.linspace(0, 1, steps).view(-1, *([1] * x.dim()))
    path = baseline + alphas * (x - baseline)               # (steps, ...)
    path.requires_grad_(True)
    out = model(path)[:, target].sum()
    grad = torch.autograd.grad(out, path)[0]
    mean_grad = grad.mean(dim=0)
    ig = (x - baseline) * mean_grad

    # the completeness check: the sum should ≈ f(x) − f(baseline)
    with torch.no_grad():
        diff = (model(x[None])[0, target] - model(baseline[None])[0, target]).item()
    return ig, {"ig_sum": ig.sum().item(), "f_diff": diff}

ig, check = integrated_gradients(model, image, target=cls)
print(check)   # {'ig_sum': 3.81, 'f_diff': 3.94}  ← close = the approximation is good enough

The completeness check is free and reveals straight away whether the number of steps is too small.

Mastery means

  • Computes gradient-based attribution
  • Knows the methods' shortcomings
  • Uses sanity checks

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences