Skip to content
AI-grafen
EUniversityRAG and information retrieval· about 60 min· fast-moving, sources checked often· verified 2026-09-20· EN

RAG — retrieval-augmented generation

Be able to build a RAG system with chunking, retrieval, the context window and citations, and to measure the answer quality.

Prerequisites

Intuition

RAG = fetch the relevant pieces of text first, then let the model answer with them in the prompt. The model does not need to «know» your documents — it reads them every time. The advantages: updatable (swap the documents, not the model), traceable (the answer can point at the source), cheaper than fine-tuning.

The chain:

  1. Chunk the documents into 200–500 tokens with overlap, keeping the metadata (title, heading, URL).
  2. Index: embeddings plus BM25.
  3. Retrieve the top k (5–10) for the question, ideally hybrid plus reranking.
  4. Prompt: «Answer only from the context. Give the source [1], [2]. If the answer is not there: say so.»
  5. Measure in two layers: was the right chunk among the top k? (retrieval) and was the answer correct and grounded? (generation).

Most RAG faults are retrieval faults. Fix those first.

Code

def chunk(text, size=400, overlap=80):
    words = text.split(); out = []
    for i in range(0, len(words), size - overlap):
        out.append(" ".join(words[i:i + size]))
    return out

def build_prompt(question, hits):
    ctx = "\n\n".join(f"[{i+1}] ({h['source']}) {h['text']}" for i, h in enumerate(hits))
    return ("Answer the question ONLY from the context below. Cite the sources as [n]. "
            "If the context is not enough, answer: 'That is not stated in the material.'\n\n"
            f"Context:\n{ctx}\n\nQuestion: {question}\nAnswer:")

def rag(question, index, llm, k=6):
    hits = index.hybrid_search(question, k=k)          # BM25 + embeddings, RRF
    answer = llm(build_prompt(question, hits))
    return {"answer": answer, "sources": [h["source"] for h in hits]}

# the eval, in two layers
retrieval_recall = mean(true_chunk in [h["id"] for h in index.hybrid_search(q, 6)] for q, true_chunk in test_cases)
grounded = mean(judge(answer, context) for ...)       # LLM-as-judge: is every claim supported by the context?

Mastery means

  • Builds a RAG chain: chunking → retrieval → a prompt with context and citations
  • Measures retrieval (recall@k) and answer quality separately
  • Identifies the most common faults: the wrong chunk, too much context, a hallucinated source

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences