Evaluating retrieval: recall@k, MRR, nDCG
Be able to build a test set and measure the retrieval quality.
Prerequisites
- DRetrieval — finding the right textrequired
- EModel evaluationrequired
Intuition
RAG has two layers that can fail. Measure them separately — otherwise you are debugging blind.
The retrieval measures:
| The measure | Asks | The formula |
|---|---|---|
| recall@k | was the right passage among the top k? | hits/questions |
| precision@k | what share of the top k were relevant? | relevant/k |
| MRR | how high up did the first right one come? | the mean of 1/rank |
| nDCG@k | rewards several relevant ones, high up | the discounted gain / the ideal |
recall@k matters most in RAG: if the right passage is not among what is sent to the model the answer can never be right. MRR and nDCG matter when the order has an effect (the model reads the first ones best).
Code
import numpy as np
def recall_at_k(results, key, k):
return float(np.mean([any(d in f for d in r[:k]) for r, f in zip(results, key)]))
def mrr(results, key):
rr = []
for r, f in zip(results, key):
rank = next((i + 1 for i, d in enumerate(r) if d in f), None)
rr.append(1 / rank if rank else 0.0)
return float(np.mean(rr))
def ndcg_at_k(results, key, k):
def dcg(rel):
return sum(r / np.log2(i + 2) for i, r in enumerate(rel))
out = []
for r, f in zip(results, key):
rel = [1.0 if d in f else 0.0 for d in r[:k]]
ideal = sorted([1.0] * min(len(f), k) + [0.0] * k, reverse=True)[:k]
out.append(dcg(rel) / dcg(ideal) if dcg(ideal) else 0.0)
return float(np.mean(out))
res = [["d3", "d1", "d9"], ["d5", "d2", "d7"]]
kee = [{"d1"}, {"d7"}]
print(recall_at_k(res, kee, 3), round(mrr(res, kee), 3), round(ndcg_at_k(res, kee, 3), 3))
# 1.0 0.417 0.565
Build the test set like this: take 50–100 real questions from the logs (or from the people who are going to use the system), find which passage actually answers them, and save the pair. Questions you think up yourself are too easy and too well formulated — that is the most common reason a RAG system looks good in testing and bad in operation.
Mastery means
- Builds a retrieval test set
- Computes recall@k, MRR and nDCG
- Separates retrieval errors from generation errors
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Discounted cumulative gain (CC BY-SA 4.0) — CC BY-SA 4.0
- arXiv — BEIR: A Heterogeneous Benchmark for Zero-shot Information Retrieval — arXiv (open access; licence per article)