Skip to content
AI-grafen
EUniversityLab· about 60 min· server sandbox

Lab: BM25 and a minimal RAG pipeline

Implement BM25 from the formula, retrieve the right passage for a question and build a prompt with citations — the whole RAG chain without an LLM.

Theory

BM25(q, d) = Σ IDF(t) · tf(t,d)(k₁+1) / (tf(t,d) + k₁(1 − b + b·|d|/avgdl)). IDF(t) = ln((N − n(t) + 0.5)/(n(t) + 0.5) + 1).

Sub-tasks

  1. tokenize + idf — tokenize(text) (lower case, a–ö letters and digits); idf(term, docs).
  2. bm25 — bm25(query, docs, k1=1.5, b=0.75) → a score per document.
  3. rag_prompt — rag_prompt(query, docs, k) → a string with the k best passages numbered [1]…[k] + the question.

Passes when: recall >= 0.8

The starter code

runs in an isolated sandbox on the server
import math
import re
from collections import Counter


def tokenize(text):
    # TODO: re.findall(r"[a-zåäö0-9]+", text.lower())
    ...


def idf(term, docs):
    # TODO: docs = lista av tokenlistor
    ...


def bm25(query, docs, k1=1.5, b=0.75):
    # TODO: returnera lista med poäng per dokument
    ...


def rag_prompt(query, docs_text, k=2):
    # TODO: rangordna docs_text (strängar) med bm25, bygg "[1] ...\n[2] ...\n\nFråga: ..."
    ...

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

recall@1 ≥ 0.8 on the 10 questions against 12 passages.

Common mistakes

  • IDF without +1 can become negative for common words.
  • Forgets length normalisation (b).
  • Does not tokenize the question the same way as the documents.