Skip to content
AI-grafen

The goal F AI engineering

Deep reinforcement learning

Policy gradient and REINFORCE, DQN with a replay buffer and a target network, actor–critic and PPO — the methods behind everything from Atari to RLHF, and the faults that stop them learning.

Knowledge nodes
52
From zero
about 43 h
Labs
6
See what you already know — no account

The diagnostic removes what you already know, so your path is usually much shorter.

What you can do afterwards

Labs along the way

You write the code. Tests you cannot see decide whether it holds up.

The whole path

Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map

AExplorer2 knowledge nodes
  1. Patterns and categories
  2. Sequences and precise instructions
BInvestigator6 knowledge nodes
  1. Data in everyday life
  2. Algorithmic thinking
  3. Rule-based systems and machine learning
  4. Source criticism and responsibility in AI use
  5. Training data, features and labels
  6. Classification: how a model sorts information
CBuilder9 knowledge nodes
  1. Functions and coordinate systems
  2. Programming logic — variables, conditions, loops
  3. Python — the basics
  4. Python — lists, loops and dictionaries
  5. Statistics — mean, median and spread
  6. Probability — the basics
  7. Linear regression: fitting a straight line to data
  8. Neural networks — the intuition
  9. Prompting — steering a language model
DAI developer19 knowledge nodes
  1. Derivatives and optimisation
  2. Conditional probability
  3. Loss functions
  4. Tokenisation
  5. Gradient descent
  6. Vectors
  7. Matrices and matrix multiplication
  8. Linear regression with several features
  9. Neural networks — the forward pass with matrices
  10. Backpropagation
  11. Embeddings — words as vectors
  12. Attention
  13. Overfitting and generalisation
  14. Partial derivatives and the gradient
  15. PyTorch — tensors and autograd
  16. Retrieval — finding the right text
  17. Training, validation and test
  18. Train a neural network in PyTorch
  19. Transformers — the architecture
EUniversity8 knowledge nodes
  1. Reinforcement learning — the basics
  2. Markov decision processes
  3. Q-learning
  4. The chain rule in several variables
  5. Language models — training and generation
  6. Model evaluation
  7. Fine-tuning language models
  8. RAG — retrieval-augmented generation
FAI engineering7 knowledge nodes
  1. Deep Q-Networks
  2. Policy gradient and REINFORCE
  3. Actor–critic and PPO
  4. Evals for language models and agents
  5. RLHF and preference learning
  6. Reward hacking
  7. DPO and direct preference optimisation
GFrontier Lab1 knowledge nodes
  1. Alignment — the problems and the methods