Skip to content
AI-grafen
EUniversityRAG and information retrieval· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Evaluating retrieval: recall@k, MRR, nDCG

Be able to build a test set and measure the retrieval quality.

Prerequisites

Intuition

RAG has two layers that can fail. Measure them separately — otherwise you are debugging blind.

The retrieval measures:

The measureAsksThe formula
recall@kwas the right passage among the top k?hits/questions
precision@kwhat share of the top k were relevant?relevant/k
MRRhow high up did the first right one come?the mean of 1/rank
nDCG@krewards several relevant ones, high upthe discounted gain / the ideal

recall@k matters most in RAG: if the right passage is not among what is sent to the model the answer can never be right. MRR and nDCG matter when the order has an effect (the model reads the first ones best).

Code

import numpy as np

def recall_at_k(results, key, k):
    return float(np.mean([any(d in f for d in r[:k]) for r, f in zip(results, key)]))

def mrr(results, key):
    rr = []
    for r, f in zip(results, key):
        rank = next((i + 1 for i, d in enumerate(r) if d in f), None)
        rr.append(1 / rank if rank else 0.0)
    return float(np.mean(rr))

def ndcg_at_k(results, key, k):
    def dcg(rel):
        return sum(r / np.log2(i + 2) for i, r in enumerate(rel))
    out = []
    for r, f in zip(results, key):
        rel = [1.0 if d in f else 0.0 for d in r[:k]]
        ideal = sorted([1.0] * min(len(f), k) + [0.0] * k, reverse=True)[:k]
        out.append(dcg(rel) / dcg(ideal) if dcg(ideal) else 0.0)
    return float(np.mean(out))

res = [["d3", "d1", "d9"], ["d5", "d2", "d7"]]
kee = [{"d1"}, {"d7"}]
print(recall_at_k(res, kee, 3), round(mrr(res, kee), 3), round(ndcg_at_k(res, kee, 3), 3))
# 1.0 0.417 0.565

Build the test set like this: take 50–100 real questions from the logs (or from the people who are going to use the system), find which passage actually answers them, and save the pair. Questions you think up yourself are too easy and too well formulated — that is the most common reason a RAG system looks good in testing and bad in operation.

Mastery means

  • Builds a retrieval test set
  • Computes recall@k, MRR and nDCG
  • Separates retrieval errors from generation errors

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences