The goal F AI engineering
Interpreting a language model
Attention patterns, the logit lens and explanations for end users — how to see what a model is doing, and where the limits of what an explanation proves lie.
- Knowledge nodes
- 66
- From zero
- about 48 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: a neural network in pure NumPy — forward, backprop, gradient checkDa sandbox · about 75 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: gradient descent from scratchDin the browser · about 50 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: scaled dot-product attention with a causal maskDa sandbox · about 60 minLab: semantic search with embeddingsDa sandbox · about 45 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer4 knowledge nodes
BInvestigator9 knowledge nodes
CBuilder18 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Statistics — mean, median and spread
- Probability — the basics
- Digital representation: text, images and audio
- Proportionality and scale
- Linear relationships in tables
- Linear regression: fitting a line to data
- Linear regression: fitting a straight line to data
- Neural networks — the intuition
- Prompting — steering a language model
- Responsible use of AI
- Images as matrices
- Turn the knobs: weights
- Why multiple layers are needed
- Training a neural network in the browser
DAI developer19 knowledge nodes
- Derivatives and optimisation
- Loss functions
- Tokenisation
- Gradient descent
- Vectors
- Matrices and matrix multiplication
- Linear regression with several features
- Logistic regression and decision trees
- Neural networks — the forward pass with matrices
- Backpropagation
- What does the network see? Images through the layers
- Embeddings — words as vectors
- Attention
- Overfitting and generalisation
- PyTorch — tensors and autograd
- Training, validation and test
- Train a neural network in PyTorch
- Transformers — the architecture
- Looking inside the network
EUniversity12 knowledge nodes
- Explainability for end users
- Linear maps
- Eigenvalues and eigenvectors
- Matrix factorisation and low-rank approximation
- Visualising attention patterns
- Singular value decomposition (SVD)
- PCA — principal component analysis
- Autoencoders
- Multi-head attention in detail
- Language models — training and generation
- Model evaluation
- Scientific method in AI