FAI engineeringRAG and information retrieval· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Evaluating RAG answers
Be able to measure grounding, correctness and citation in RAG answers.
Prerequisites
- ERAG — retrieval-augmented generationrequired
- FEvals for language models and agentsrequired
Intuition
A RAG answer can fail in four different ways, and they require different measures:
| The fault | The measure | The remedy |
|---|---|---|
| The right passage was never fetched | recall@k | the chunking, the embeddings, the hybrid, reranking |
| The right passage was fetched but not used | context precision / utilisation | reranking, fewer chunks, the prompt |
| The answer is not supported by the passages | grounding | a citation requirement, a declining instruction, a verification step |
| The answer is grounded but does not answer the question | answer relevance | the prompt, question understanding |
Measure all four. A single «quality measure» hides which of them is the problem — and therefore which remedy helps.
Code
import json
GROUNDING = """Split the ANSWER into individual claims. For each: is it supported by the CONTEXT?
Answer in JSON: {"claims": [{"text": "...", "supported": true|false}]}"""
RELEVANCE = """Does the ANSWER answer the question? Judge the relevance only, not the correctness.
Answer in JSON: {"relevance": 0|1, "motivation": "..."}"""
def evaluate_rag(llm, cases, chain):
out = []
for f in cases: # f: {question, key_chunk_ids, key_answer}
hits, answer = chain(f["question"])
fetched = [h["id"] for h in hits]
context = "\n".join(h["text"] for h in hits)
g = json.loads(llm(f"{GROUNDING}\nCONTEXT:\n{context}\nANSWER:\n{answer}", temperature=0, json_mode=True))
r = json.loads(llm(f"{RELEVANCE}\nQUESTION: {f['question']}\nANSWER: {answer}", temperature=0, json_mode=True))
c = g["claims"]
out.append({
"recall": float(any(x in fetched for x in f["key_chunk_ids"])),
"grounding": sum(x["supported"] for x in c) / max(len(c), 1),
"relevance": float(r["relevance"]),
"correct": float(judge_against_key(answer, f["key_answer"])),
})
return {k: round(sum(r[k] for r in out) / len(out), 3) for k in out[0]}
# {'recall': 0.91, 'grounding': 0.74, 'relevance': 0.96, 'correct': 0.68}
# → the retrieval is good, the grounding is the problem: the model is filling in from prior knowledge
The test set decides everything: 50–100 real questions from the logs, with the right passage pointed out by a human being. Questions you think up yourself are well formulated in a way real questions are not — and give a system that looks good in testing and worse in operation.
Mastery means
- Measures grounding, correctness and citation separately
- Builds a test set out of real questions
- Locates the fault in the retrieval or the generation
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — RAGAS: Automated Evaluation of RAG — arXiv (open access; licence per article)
- arXiv — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — arXiv (open access; licence per article)