Hallucinations — causes and countermeasures
Be able to explain why models make things up and which methods (grounding, calibration, declining) reduce it.
Prerequisites
Intuition
A model hallucinates when it produces something that sounds plausible but is not true. It is not a bug that can be fixed away — it follows from how the model works.
Why it happens:
- The training objective rewards likely text, not true text. An invented year in the right place in a sentence is likely.
- The model has no internal marking of what it «knows». It is all weights.
- Preference training (RLHF) rewards helpful, confident answers — and rarely punishes «I don't know» hard enough.
- Rare facts have a weak signal in the training data; the model interpolates.
Where it is worst: exact numbers, years, quotations, citations, people's names, rare topics, and anything that happened after the training data ended.
Formal
The countermeasures, in order of effectiveness:
| Method | What it does | The limitation |
|---|---|---|
| Grounding (RAG) | give the model the source text in the prompt | only helps if the right text is retrieved |
| A citation requirement | every claim has to point at a retrieved passage | the model can cite the wrong passage |
| A declining instruction | «answer 'not stated' if the material is missing» | lowers recall, has to be trained or prompted in |
| Self-consistency | sample n answers, take the majority | expensive, helps most on reasoning |
| A verification step | a second model checks against the source | double the cost |
| Calibration | the model's confidence should match its accuracy | models are often overconfident |
Measure it: take 100 questions with known answers, and let a judge (an LLM or a human) classify every claim as supported by the source / contradicted / not verifiable. Report the share of grounded claims — that is the hallucination rate, and it can be tracked between versions.
Code
GROUNDING_PROMPT = """Answer ONLY from the SOURCES. Every claim must be followed by [n] pointing at the source.
If the sources are not enough: answer exactly "That is not stated in the material." Never guess.
SOURCES:
{sources}
QUESTION: {question}"""
JUDGE = """For every claim in the ANSWER, decide whether it is supported by the SOURCES.
Answer as JSON: {"claims": [{"text": "...", "supported": true/false}]}"""
def groundedness(llm, answer, sources):
d = json.loads(llm(JUDGE + f"\nSOURCES:\n{sources}\nANSWER:\n{answer}", temperature=0, json_mode=True))
c = d["claims"]
return sum(x["supported"] for x in c) / max(len(c), 1)
Important to know: grounding does not remove the problem, it moves it. The model can still summarise wrongly, conflate two sources or cite a passage that does not say what it claims. Which is why you measure the groundedness — rather than assuming it.
Mastery means
- Explains why models make things up
- Chooses the countermeasure: grounding, calibration, declining
- Measures the hallucination rate
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Survey of Hallucination in Natural Language Generation — arXiv (open access; licence per article)
- arXiv — RAGAS: Automated Evaluation of RAG — arXiv (open access; licence per article)