Final project: build a traceable RAG pipeline
Build a small document assistant that chunks documents without losing their source, retrieves relevant passages, returns a verbatim answer with a checkable citation, and abstains when evidence is missing. After passing the code project, you can test how a real AI model responds to selected sources.
Theory
A reliable RAG pipeline can be evaluated in four stages: documents → source-labelled chunks → relevant hit → answer with checkable evidence. Measure retrieval separately from answer quality. Abstain when evidence is missing.
Sub-tasks
- Keep source labels when chunking —
chunk_documents(documents, max_words)returns chunks withsource_idandtext. Preserve every word in order, never mix sources, and keep each chunk withinmax_wordswords. - Retrieve relevant evidence —
retrieve(query, chunks, k)returns at most k chunks ranked by relevance. Ignore common question words; return an empty list when content words have no match.evaluate_retrieval(cases, chunks, k)measures the share of queries where the correct source is among the hits. - Answer with a checkable citation —
answer_with_evidence(query, hits)returns{answer, citations}. The answer must be a verbatim sentence from a hit. Each citation is{source_id, quote}with the same exact sentence. If nothing relevant is found, use the specified Swedish abstention string and an empty citations list.
Passes when: recall >= 0.8
The starter code
runs in an isolated sandbox on the server"""RAG-slutprov. Ändra funktionerna nedan; testerna bedömer beteende på nya dokument."""
ABSTAIN = "Jag hittar inget stöd i underlaget."
def chunk_documents(documents, max_words=30):
"""documents: [{id, text}] -> [{source_id, text}], i ursprunglig ordning."""
# TODO: dela varje dokuments ord i stycken utan att blanda källor.
raise NotImplementedError
def retrieve(query, chunks, k=2):
"""Returnera de mest relevanta källmärkta styckena; [] om inget innehållsord matchar."""
# TODO: rangordna med en reproducerbar lexikal poäng; undvik träffar på bara frågeord.
raise NotImplementedError
def answer_with_evidence(query, hits):
"""Returnera {answer: str, citations: [{source_id, quote}]} eller ett belagt avstående."""
# TODO: välj ett relevant meningsutdrag från träffarna och citera exakt den källan.
raise NotImplementedError
def evaluate_retrieval(cases, chunks, k=2):
"""cases: [{query, source_id}] -> recall@k som flyttal mellan 0 och 1."""
# TODO: använd retrieve och räkna hur ofta rätt källa finns bland de k första träffarna.
raise NotImplementedError
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
All tests and recall@2 ≥ 0.8. The server also runs hidden cases with different documents, sources, and queries. The course is complete only after a verified full run of this project and all knowledge steps.
Common mistakes
Losing source_id during chunking; returning the first chunks for every query; inventing a citation; guessing without evidence; running only one subtask instead of All.