Skip to content
AI-grafen

The goal F AI engineering

Run models more cheaply: quantisation

Shrink a model to a quarter of its size, measure what is lost, and choose the right format for your runtime.

Knowledge nodes
59
From zero
about 46 h
Labs
6
See what you already know — no account

The diagnostic removes what you already know, so your path is usually much shorter.

What you can do afterwards

Labs along the way

You write the code. Tests you cannot see decide whether it holds up.

The whole path

Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map

AExplorer4 knowledge nodes
  1. Patterns and categories
  2. Sequences and precise instructions
  3. Integer arithmetic and order of operations
  4. Comparing and sorting values
BInvestigator8 knowledge nodes
  1. Data in everyday life
  2. Algorithmic thinking
  3. Rule-based systems and machine learning
  4. Coordinates: positions on a grid
  5. Training data, features and labels
  6. Classification: how a model sorts information
  7. Generative AI: how it creates content
  8. Language models and probabilities
CBuilder11 knowledge nodes
  1. Functions and coordinate systems
  2. Programming logic — variables, conditions, loops
  3. Python — the basics
  4. Python — lists, loops and dictionaries
  5. Statistics — mean, median and spread
  6. Probability — the basics
  7. Linear regression: fitting a straight line to data
  8. Neural networks — the intuition
  9. Attention: how words influence each other
  10. Words as points: similar words close together
  11. Why does it take time for the AI to answer?
DAI developer20 knowledge nodes
  1. Derivatives and optimisation
  2. Time complexity and big-O notation
  3. Python — functions, scope and exceptions
  4. Python — modules, packages and virtual environments
  5. Loss functions
  6. The terminal and the shell
  7. Tokenisation
  8. Gradient descent
  9. The transformer — an overview without formulas
  10. How a language model is trained — at upper-secondary level
  11. Vectors
  12. Matrices and matrix multiplication
  13. Neural networks — the forward pass with matrices
  14. Backpropagation
  15. Embeddings — words as vectors
  16. Attention
  17. PyTorch — tensors and autograd
  18. Train a neural network in PyTorch
  19. Transformers — the architecture
  20. A small model or a large one?
EUniversity9 knowledge nodes
  1. Docker — containers
  2. Context length and the quadratic cost
  3. Running models locally: llama.cpp, vLLM, Ollama
  4. Multi-head attention in detail
  5. The KV cache
  6. Latency, throughput and batching in inference
  7. Cost modelling for LLM systems
  8. Language models — training and generation
  9. Sampling: temperature, top-k, top-p, beam
FAI engineering6 knowledge nodes
  1. Quantisation
  2. Quantisation in practice: GPTQ, AWQ, GGUF
  3. Inference optimisation
  4. Prompt caching and prefix sharing
  5. Multi-query and grouped-query attention
  6. Speculative decoding
GFrontier Lab1 knowledge nodes
  1. Pruning and sparsity