Why did it answer that?
Be able to ask for an explanation and judge whether it is credible.
Prerequisites
Intuition
The model says «cat» about your picture. Why?
There are two kinds of answer:
- What the model looked at. Tools can colour in the pixels that affected the decision most. If the ears light up that is reasonable. If the sofa in the background lights up the model has found a shortcut.
- An explanation in words. A language model can write «I answered that because …». It sounds convincing — but it is a post hoc construction. The model has no insight into its own computations; it generates a plausible explanation.
The difference matters: the first is a measurement, the second is a guess.
Intuition
Test whether the explanation holds. If the model says «I saw the pointed ears»:
- Cover the ears in the picture and ask again. Does the answer change? Then the explanation was right.
- Cover the background instead. Does the answer change then? Then it was the background that decided, whatever the model claimed.
This is called an intervention and is the only way of knowing. An explanation that cannot be tested is a story.
The same goes for text models: ask «why did you answer that?» and you often get a neat justification — even for an answer that was wrong. Ask for the source instead and check it.
Mastery means
- Asks for an explanation of a model answer
- Judges whether the explanation is credible
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Grad-CAM — arXiv (open access; licence per article)
- arXiv — Language Models Don't Always Say What They Think — arXiv (open access; licence per article)