EUniversityLab· about 60 min· server sandbox
Lab: BM25 and a minimal RAG pipeline
Implement BM25 from the formula, retrieve the right passage for a question and build a prompt with citations — the whole RAG chain without an LLM.
Teaches: BM25 and keyword search
Requires: Text preprocessingRetrieval — finding the right text
Theory
BM25(q, d) = Σ IDF(t) · tf(t,d)(k₁+1) / (tf(t,d) + k₁(1 − b + b·|d|/avgdl)). IDF(t) = ln((N − n(t) + 0.5)/(n(t) + 0.5) + 1).
Sub-tasks
- tokenize + idf —
tokenize(text)(lower case, a–ö letters and digits);idf(term, docs). - bm25 —
bm25(query, docs, k1=1.5, b=0.75)→ a score per document. - rag_prompt —
rag_prompt(query, docs, k)→ a string with the k best passages numbered [1]…[k] + the question.
Passes when: recall >= 0.8
The starter code
runs in an isolated sandbox on the serverimport math
import re
from collections import Counter
def tokenize(text):
# TODO: re.findall(r"[a-zåäö0-9]+", text.lower())
...
def idf(term, docs):
# TODO: docs = lista av tokenlistor
...
def bm25(query, docs, k1=1.5, b=0.75):
# TODO: returnera lista med poäng per dokument
...
def rag_prompt(query, docs_text, k=2):
# TODO: rangordna docs_text (strängar) med bm25, bygg "[1] ...\n[2] ...\n\nFråga: ..."
...
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
recall@1 ≥ 0.8 on the 10 questions against 12 passages.
Common mistakes
- IDF without +1 can become negative for common words.
- Forgets length normalisation (b).
- Does not tokenize the question the same way as the documents.