EUniversityRAG and information retrieval· about 60 min· fast-moving, sources checked often· verified 2026-09-20· EN
RAG — retrieval-augmented generation
Be able to build a RAG system with chunking, retrieval, the context window and citations, and to measure the answer quality.
Prerequisites
Intuition
RAG = fetch the relevant pieces of text first, then let the model answer with them in the prompt. The model does not need to «know» your documents — it reads them every time. The advantages: updatable (swap the documents, not the model), traceable (the answer can point at the source), cheaper than fine-tuning.
The chain:
- Chunk the documents into 200–500 tokens with overlap, keeping the metadata (title, heading, URL).
- Index: embeddings plus BM25.
- Retrieve the top k (5–10) for the question, ideally hybrid plus reranking.
- Prompt: «Answer only from the context. Give the source [1], [2]. If the answer is not there: say so.»
- Measure in two layers: was the right chunk among the top k? (retrieval) and was the answer correct and grounded? (generation).
Most RAG faults are retrieval faults. Fix those first.
Code
def chunk(text, size=400, overlap=80):
words = text.split(); out = []
for i in range(0, len(words), size - overlap):
out.append(" ".join(words[i:i + size]))
return out
def build_prompt(question, hits):
ctx = "\n\n".join(f"[{i+1}] ({h['source']}) {h['text']}" for i, h in enumerate(hits))
return ("Answer the question ONLY from the context below. Cite the sources as [n]. "
"If the context is not enough, answer: 'That is not stated in the material.'\n\n"
f"Context:\n{ctx}\n\nQuestion: {question}\nAnswer:")
def rag(question, index, llm, k=6):
hits = index.hybrid_search(question, k=k) # BM25 + embeddings, RRF
answer = llm(build_prompt(question, hits))
return {"answer": answer, "sources": [h["source"] for h in hits]}
# the eval, in two layers
retrieval_recall = mean(true_chunk in [h["id"] for h in index.hybrid_search(q, 6)] for q, true_chunk in test_cases)
grounded = mean(judge(answer, context) for ...) # LLM-as-judge: is every claim supported by the context?
Mastery means
- Builds a RAG chain: chunking → retrieval → a prompt with context and citations
- Measures retrieval (recall@k) and answer quality separately
- Identifies the most common faults: the wrong chunk, too much context, a hallucinated source
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — arXiv (open access; licence per article)
- arXiv — RAGAS: Automated Evaluation of RAG — arXiv (open access; licence per article)
Leads to
Part of the goals (13)
- Build a RAG system you can trust
- Build an AI service that survives production
- Multimodal systems
- AI in production
- AI safety in practice
- Frontier Lab — an independent research project
- Fine-tune and run your own models
- Build a memory system for an agent
- Build an agent you can trust
- Evals in practice
- An AI service in operation
- Reproduce a paper
- Deep reinforcement learning