The goal F AI engineering
AI safety in practice
Red teaming, jailbreaks, adversarial examples, preference learning and reward hacking — what can be defended against and what cannot.
- Knowledge nodes
- 65
- From zero
- about 57 h
- Labs
- 6
See what you already know — no account
The diagnostic removes what you already know, so your path is usually much shorter.
What you can do afterwards
- FAdversarial examples
- FRLHF and preference learning
- GAlignment — the problems and the methods
- FReward hacking
- FDPO and direct preference optimisation
- GConstitutional AI and rule-based alignment
- FSpecification problems and proxy objectives
- FAI safety and red teaming
- FJailbreaks and guard rails
- FData poisoning and backdoors
Labs along the way
You write the code. Tests you cannot see decide whether it holds up.
Lab: a neural network in pure NumPy — forward, backprop, gradient checkDa sandbox · about 75 minLab: dot product, norm and cosine similarityDin the browser · about 40 minLab: gradient descent from scratchDin the browser · about 50 minLab: matrix multiplication and one layer of a neural networkDin the browser · about 45 minLab: scaled dot-product attention with a causal maskDa sandbox · about 60 minLab: semantic search with embeddingsDa sandbox · about 45 min
The whole path
Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map
AExplorer2 knowledge nodes
BInvestigator6 knowledge nodes
CBuilder12 knowledge nodes
- Functions and coordinate systems
- Programming logic — variables, conditions, loops
- Python — the basics
- Python — lists, loops and dictionaries
- Python — strings and text processing
- Python — files, CSV and JSON
- Statistics — mean, median and spread
- Probability — the basics
- Linear regression: fitting a straight line to data
- Neural networks — the intuition
- Prompting — steering a language model
- Responsible use of AI
DAI developer22 knowledge nodes
- Derivatives and optimisation
- APIs and HTTP
- Regular expressions
- Licences and open data
- Loss functions
- Tokenisation
- Gradient descent
- The context window, system prompts and few-shot
- Vectors
- Matrices and matrix multiplication
- Linear regression with several features
- Neural networks — the forward pass with matrices
- Backpropagation
- Embeddings — words as vectors
- Attention
- Overfitting and generalisation
- Partial derivatives and the gradient
- PyTorch — tensors and autograd
- Retrieval — finding the right text
- Training, validation and test
- Train a neural network in PyTorch
- Transformers — the architecture
EUniversity10 knowledge nodes
FAI engineering11 knowledge nodes
- Adversarial examples
- Dataset design for fine-tuning
- Evals for language models and agents
- RLHF and preference learning
- Reward hacking
- DPO and direct preference optimisation
- Specification problems and proxy objectives
- AI safety and red teaming
- Jailbreaks and guard rails
- Training data for language models: filtering and dedup
- Data poisoning and backdoors