Skip to content
AI-grafen
DAI developerRAG and information retrieval· about 45 min· evolving, reviewed regularly· verified 2026-09-20· EN

Retrieval — finding the right text

Be able to build semantic search with embeddings, combine it with keyword search (BM25) and evaluate it with recall@k.

Prerequisites

Intuition

Retrieval = finding the pieces of text that best answer a question among thousands.

  • Keyword search (BM25): counts word overlap, weighted by how uncommon the words are. Fast, exact on names and codes, blind to synonyms.
  • Semantic search: the query and every piece of text become an embedding (a vector); the closest in angle (cosine similarity) wins. Finds «car» when you ask about «vehicle», sometimes misses exact terms.
  • Hybrid: run both and merge the rankings (with RRF, say). Almost always the best.

Measure with recall@k: for every test question with a known correct passage — was it among the first k?

Code

import numpy as np

def cos(a, b):
    return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))

def semantic_topk(q_emb, doc_embs, k=5):
    s = [cos(q_emb, d) for d in doc_embs]
    return list(np.argsort(s)[::-1][:k])

def rrf(rankings, k=60):
    scores = {}
    for r in rankings:
        for pos, doc in enumerate(r):
            scores[doc] = scores.get(doc, 0) + 1 / (k + pos + 1)
    return sorted(scores, key=scores.get, reverse=True)

def recall_at_k(results, truth, k=10):
    return np.mean([t in r[:k] for r, t in zip(results, truth)])

BM25 is in rank_bm25; the embeddings come from an embedding model (via an API or sentence-transformers). Chunk the documents into 200–500 tokens with overlap before you embed them.

Mastery means

  • Builds semantic search with embeddings and cosine similarity
  • Combines it with BM25 (hybrid)
  • Measures recall@k

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences