The goal F AI engineering
Evals in practice
LLM as judge, regression tests, statistical significance and how to value negative results — measurement that holds between releases.
- Knowledge nodes
- 62
- From zero
- about 54 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: build an eval harnessEa sandbox · about 60 minLab: a neural network in pure NumPy — forward, backprop, gradient checkDa sandbox · about 75 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: gradient descent from scratchDin the browser · about 50 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: scaled dot-product attention with a causal maskDa sandbox · about 60 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer2 knowledge nodes
BInvestigator6 knowledge nodes
CBuilder9 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Statistics — mean, median and spread
- Probability — the basics
- Linear regression: fitting a straight line to data
- Neural networks — the intuition
- Prompting — steering a language model
DAI developer23 knowledge nodes
- Derivatives and optimisation
- Python — functions, scope and exceptions
- Correlation and causation
- Probability distributions
- Loss functions
- The normal distribution and standardisation
- Samples and uncertainty
- Testing with pytest
- Tokenisation
- Gradient descent
- Vectors
- Matrices and matrix multiplication
- Linear regression with several features
- Neural networks — the forward pass with matrices
- Backpropagation
- Embeddings — words as vectors
- Attention
- Overfitting and generalisation
- PyTorch — tensors and autograd
- Retrieval — finding the right text
- Training, validation and test
- Train a neural network in PyTorch
- Transformers — the architecture
EUniversity10 knowledge nodes
FAI engineering12 knowledge nodes
- Causal inference
- Multiple comparisons
- Negative results and publication bias
- Human evaluation and annotator agreement
- Evals for language models and agents
- An LLM as a judge
- Regression tests for models
- Evals for code generation
- Statistical significance in evals
- Evaluating an AI tutor
- Reading and analysing research papers
- Keeping up with the research front