The goal E University
Build a RAG system you can trust
Chunking, retrieval evaluation, citation and grounding — the whole chain from document to answer with traceable sources.
- Knowledge nodes
- 77
- From zero
- about 58 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
- DBM25 and keyword search
- DBuild a small search engine
- EStoring and finding vectors efficiently
- EChunking documents
- EHybrid search and RRF
- FReranking with cross-encoders
- FChoosing and fine-tuning embedding models
- ERAG — retrieval-augmented generation
- ECitation and grounding
- FKnowledge graphs and graph RAG
- FLong context versus retrieval
- FEvaluating RAG answers
- EEvaluating retrieval: recall@k, MRR, nDCG
- EIndex updating and versioning
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: BM25 and a minimal RAG pipelineEa sandbox · about 60 minLab: a neural network in pure NumPy — forward, backprop, gradient checkDa sandbox · about 75 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: gradient descent from scratchDin the browser · about 50 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: scaled dot-product attention with a causal maskDa sandbox · about 60 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer3 knowledge nodes
BInvestigator10 knowledge nodes
- Data in everyday life
- Algorithmic thinking
- Rule-based systems and machine learning
- Source criticism and responsibility in AI use
- Training data, features and labels
- Classification: how a model sorts information
- The technology behind the web: how a page is fetched
- Generative AI: how it creates content
- Language models and probabilities
- AI's invented answers
CBuilder15 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Python — strings and text processing
- Python — files, CSV and JSON
- Search strategies: linear search vs binary search
- Statistics — mean, median and spread
- Probability — the basics
- Linear regression: fitting a straight line to data
- Neural networks — the intuition
- Prompting — steering a language model
- Fundamentals of search engines: indexing and ranking
- Asking effective questions to AI
- RAG: letting AI answer based on your own documents
DAI developer29 knowledge nodes
- Discrete mathematics: graphs and relations
- Derivatives and optimisation
- Time complexity and big-O notation
- Data structures: lists, stacks, queues, hash tables
- Python — functions, scope and exceptions
- Python — modules, packages and virtual environments
- Git — version control
- Loss functions
- Tokenisation
- Text preprocessing
- BM25 and keyword search
- Gradient descent
- Build a small search engine
- Vectors
- Matrices and matrix multiplication
- Linear regression with several features
- Neural networks — the forward pass with matrices
- Backpropagation
- Embeddings — words as vectors
- Attention
- NumPy — arrays and vectorisation
- Overfitting and generalisation
- Pandas — tables in Python
- PyTorch — tensors and autograd
- Retrieval — finding the right text
- SQL — the basics
- Training, validation and test
- Train a neural network in PyTorch
- Transformers — the architecture
EUniversity14 knowledge nodes
- Storing and finding vectors efficiently
- Context length and the quadratic cost
- Chunking documents
- Hybrid search and RRF
- Data pipelines and ETL
- Data versioning
- Language models — training and generation
- Model evaluation
- Fine-tuning language models
- RAG — retrieval-augmented generation
- Citation and grounding
- Evaluating retrieval: recall@k, MRR, nDCG
- Vector databases and indexing
- Index updating and versioning