Skip to content
AI-grafen
FAI engineeringRAG and information retrieval· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Evaluating RAG answers

Be able to measure grounding, correctness and citation in RAG answers.

Prerequisites

Intuition

A RAG answer can fail in four different ways, and they require different measures:

The faultThe measureThe remedy
The right passage was never fetchedrecall@kthe chunking, the embeddings, the hybrid, reranking
The right passage was fetched but not usedcontext precision / utilisationreranking, fewer chunks, the prompt
The answer is not supported by the passagesgroundinga citation requirement, a declining instruction, a verification step
The answer is grounded but does not answer the questionanswer relevancethe prompt, question understanding

Measure all four. A single «quality measure» hides which of them is the problem — and therefore which remedy helps.

Code

import json

GROUNDING = """Split the ANSWER into individual claims. For each: is it supported by the CONTEXT?
Answer in JSON: {"claims": [{"text": "...", "supported": true|false}]}"""

RELEVANCE = """Does the ANSWER answer the question? Judge the relevance only, not the correctness.
Answer in JSON: {"relevance": 0|1, "motivation": "..."}"""

def evaluate_rag(llm, cases, chain):
    out = []
    for f in cases:                                  # f: {question, key_chunk_ids, key_answer}
        hits, answer = chain(f["question"])
        fetched = [h["id"] for h in hits]
        context = "\n".join(h["text"] for h in hits)

        g = json.loads(llm(f"{GROUNDING}\nCONTEXT:\n{context}\nANSWER:\n{answer}", temperature=0, json_mode=True))
        r = json.loads(llm(f"{RELEVANCE}\nQUESTION: {f['question']}\nANSWER: {answer}", temperature=0, json_mode=True))
        c = g["claims"]

        out.append({
            "recall": float(any(x in fetched for x in f["key_chunk_ids"])),
            "grounding": sum(x["supported"] for x in c) / max(len(c), 1),
            "relevance": float(r["relevance"]),
            "correct": float(judge_against_key(answer, f["key_answer"])),
        })
    return {k: round(sum(r[k] for r in out) / len(out), 3) for k in out[0]}

# {'recall': 0.91, 'grounding': 0.74, 'relevance': 0.96, 'correct': 0.68}
# → the retrieval is good, the grounding is the problem: the model is filling in from prior knowledge

The test set decides everything: 50–100 real questions from the logs, with the right passage pointed out by a human being. Questions you think up yourself are well formulated in a way real questions are not — and give a system that looks good in testing and worse in operation.

Mastery means

  • Measures grounding, correctness and citation separately
  • Builds a test set out of real questions
  • Locates the fault in the retrieval or the generation

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences