Skip to content
AI-grafen

The goal E University

Build a transformer from scratch

Implement self-attention, multi-head attention and positional encoding yourself, and put them together into a working small transformer.

Knowledge nodes
49
From zero
about 40 h
Labs
6
See what you already know — no account

The diagnostic removes what you already know, so your path is usually much shorter.

What you can do afterwards

Labs along the way

You write the code. Tests you cannot see decide whether it holds up.

The whole path

Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map

AExplorer2 knowledge nodes
  1. Patterns and categories
  2. Sequences and precise instructions
BInvestigator5 knowledge nodes
  1. Data in everyday life
  2. Algorithmic thinking
  3. Rule-based systems and machine learning
  4. Training data, features and labels
  5. Classification: how a model sorts information
CBuilder9 knowledge nodes
  1. Functions and coordinate systems
  2. Distance calculation and Pythagoras' theorem
  3. Programming logic — variables, conditions, loops
  4. Python — the basics
  5. Python — lists, loops and dictionaries
  6. Statistics — mean, median and spread
  7. Probability — the basics
  8. Linear regression: fitting a straight line to data
  9. Neural networks — the intuition
DAI developer19 knowledge nodes
  1. Derivatives and optimisation
  2. Time complexity and big-O notation
  3. Python — functions, scope and exceptions
  4. Python — classes and objects
  5. Probability distributions
  6. The normal distribution and standardisation
  7. Tokenisation
  8. Gradient descent
  9. Trigonometry — the basics
  10. Vectors
  11. Matrices and matrix multiplication
  12. Neural networks — the forward pass with matrices
  13. Activation functions
  14. Backpropagation
  15. Embeddings — words as vectors
  16. Attention
  17. PyTorch — tensors and autograd
  18. Train a neural network in PyTorch
  19. Transformers — the architecture
EUniversity10 knowledge nodes
  1. Context length and the quadratic cost
  2. Parallelism and why GPUs
  3. How autograd works inside
  4. Build a mini autograd
  5. Batch and layer normalisation
  6. LayerNorm, RMSNorm and pre-/post-norm
  7. The MLP block: GELU, SwiGLU
  8. Positional encoding and RoPE
  9. Multi-head attention in detail
  10. Build a small GPT from scratch
FAI engineering4 knowledge nodes
  1. The memory hierarchy and bandwidth
  2. Flash attention and IO-aware attention
  3. Mixture of Experts
  4. RoPE in detail