The goal F AI engineering
Run models more cheaply: quantisation
Shrink a model to a quarter of its size, measure what is lost, and choose the right format for your runtime.
- Knowledge nodes
- 59
- From zero
- about 46 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
- EContext length and the quadratic cost
- FQuantisation
- GPruning and sparsity
- FQuantisation in practice: GPTQ, AWQ, GGUF
- ERunning models locally: llama.cpp, vLLM, Ollama
- FInference optimisation
- FPrompt caching and prefix sharing
- FMulti-query and grouped-query attention
- ELatency, throughput and batching in inference
- ECost modelling for LLM systems
- DA small model or a large one?
- FSpeculative decoding
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: a neural network in pure NumPy — forward, backprop, gradient checkDa sandbox · about 75 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: gradient descent from scratchDin the browser · about 50 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: scaled dot-product attention with a causal maskDa sandbox · about 60 minLab: semantic search with embeddingsDa sandbox · about 45 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer4 knowledge nodes
BInvestigator8 knowledge nodes
CBuilder11 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Statistics — mean, median and spread
- Probability — the basics
- Linear regression: fitting a straight line to data
- Neural networks — the intuition
- Attention: how words influence each other
- Words as points: similar words close together
- Why does it take time for the AI to answer?
DAI developer20 knowledge nodes
- Derivatives and optimisation
- Time complexity and big-O notation
- Python — functions, scope and exceptions
- Python — modules, packages and virtual environments
- Loss functions
- The terminal and the shell
- Tokenisation
- Gradient descent
- The transformer — an overview without formulas
- How a language model is trained — at upper-secondary level
- Vectors
- Matrices and matrix multiplication
- Neural networks — the forward pass with matrices
- Backpropagation
- Embeddings — words as vectors
- Attention
- PyTorch — tensors and autograd
- Train a neural network in PyTorch
- Transformers — the architecture
- A small model or a large one?
EUniversity9 knowledge nodes
- Docker — containers
- Context length and the quadratic cost
- Running models locally: llama.cpp, vLLM, Ollama
- Multi-head attention in detail
- The KV cache
- Latency, throughput and batching in inference
- Cost modelling for LLM systems
- Language models — training and generation
- Sampling: temperature, top-k, top-p, beam