Skip to content
AI-grafen
EUniversityProject· about 75 min· server sandbox

Final project: build a traceable RAG pipeline

Build a small document assistant that chunks documents without losing their source, retrieves relevant passages, returns a verbatim answer with a checkable citation, and abstains when evidence is missing. After passing the code project, you can test how a real AI model responds to selected sources.

Theory

A reliable RAG pipeline can be evaluated in four stages: documents → source-labelled chunks → relevant hit → answer with checkable evidence. Measure retrieval separately from answer quality. Abstain when evidence is missing.

Sub-tasks

  1. Keep source labels when chunking — chunk_documents(documents, max_words) returns chunks with source_id and text. Preserve every word in order, never mix sources, and keep each chunk within max_words words.
  2. Retrieve relevant evidence — retrieve(query, chunks, k) returns at most k chunks ranked by relevance. Ignore common question words; return an empty list when content words have no match. evaluate_retrieval(cases, chunks, k) measures the share of queries where the correct source is among the hits.
  3. Answer with a checkable citation — answer_with_evidence(query, hits) returns {answer, citations}. The answer must be a verbatim sentence from a hit. Each citation is {source_id, quote} with the same exact sentence. If nothing relevant is found, use the specified Swedish abstention string and an empty citations list.

Passes when: recall >= 0.8

The starter code

runs in an isolated sandbox on the server
"""RAG-slutprov. Ändra funktionerna nedan; testerna bedömer beteende på nya dokument."""

ABSTAIN = "Jag hittar inget stöd i underlaget."


def chunk_documents(documents, max_words=30):
    """documents: [{id, text}] -> [{source_id, text}], i ursprunglig ordning."""
    # TODO: dela varje dokuments ord i stycken utan att blanda källor.
    raise NotImplementedError


def retrieve(query, chunks, k=2):
    """Returnera de mest relevanta källmärkta styckena; [] om inget innehållsord matchar."""
    # TODO: rangordna med en reproducerbar lexikal poäng; undvik träffar på bara frågeord.
    raise NotImplementedError


def answer_with_evidence(query, hits):
    """Returnera {answer: str, citations: [{source_id, quote}]} eller ett belagt avstående."""
    # TODO: välj ett relevant meningsutdrag från träffarna och citera exakt den källan.
    raise NotImplementedError


def evaluate_retrieval(cases, chunks, k=2):
    """cases: [{query, source_id}] -> recall@k som flyttal mellan 0 och 1."""
    # TODO: använd retrieve och räkna hur ofta rätt källa finns bland de k första träffarna.
    raise NotImplementedError

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

All tests and recall@2 ≥ 0.8. The server also runs hidden cases with different documents, sources, and queries. The course is complete only after a verified full run of this project and all knowledge steps.

Common mistakes

Losing source_id during chunking; returning the first chunks for every query; inventing a citation; guessing without evidence; running only one subtask instead of All.