Skip to content
AI-grafen

The goal F AI engineering

Evals in practice

LLM as judge, regression tests, statistical significance and how to value negative results — measurement that holds between releases.

Knowledge nodes
62
From zero
about 54 h
Labs
6
See what you already know — no account

The diagnostic removes what you already know, so your path is usually much shorter.

What you can do afterwards

Labs along the way

You write the code. Tests you cannot see decide whether it holds up.

The whole path

Everything the goal builds on, grouped by level and in the order it builds on itself. Show on the map

AExplorer2 knowledge nodes
  1. Patterns and categories
  2. Sequences and precise instructions
BInvestigator6 knowledge nodes
  1. Data in everyday life
  2. Algorithmic thinking
  3. Rule-based systems and machine learning
  4. Source criticism and responsibility in AI use
  5. Training data, features and labels
  6. Classification: how a model sorts information
CBuilder9 knowledge nodes
  1. Functions and coordinate systems
  2. Programming logic — variables, conditions, loops
  3. Python — the basics
  4. Python — lists, loops and dictionaries
  5. Statistics — mean, median and spread
  6. Probability — the basics
  7. Linear regression: fitting a straight line to data
  8. Neural networks — the intuition
  9. Prompting — steering a language model
DAI developer23 knowledge nodes
  1. Derivatives and optimisation
  2. Python — functions, scope and exceptions
  3. Correlation and causation
  4. Probability distributions
  5. Loss functions
  6. The normal distribution and standardisation
  7. Samples and uncertainty
  8. Testing with pytest
  9. Tokenisation
  10. Gradient descent
  11. Vectors
  12. Matrices and matrix multiplication
  13. Linear regression with several features
  14. Neural networks — the forward pass with matrices
  15. Backpropagation
  16. Embeddings — words as vectors
  17. Attention
  18. Overfitting and generalisation
  19. PyTorch — tensors and autograd
  20. Retrieval — finding the right text
  21. Training, validation and test
  22. Train a neural network in PyTorch
  23. Transformers — the architecture
EUniversity10 knowledge nodes
  1. Confidence intervals
  2. Hypothesis testing and p-values
  3. A/B tests and experiment design
  4. Annotation and labelling of data
  5. Language models — training and generation
  6. Model evaluation
  7. Build an eval harness
  8. RAG — retrieval-augmented generation
  9. Language models for code
  10. Scientific method in AI
FAI engineering12 knowledge nodes
  1. Causal inference
  2. Multiple comparisons
  3. Negative results and publication bias
  4. Human evaluation and annotator agreement
  5. Evals for language models and agents
  6. An LLM as a judge
  7. Regression tests for models
  8. Evals for code generation
  9. Statistical significance in evals
  10. Evaluating an AI tutor
  11. Reading and analysing research papers
  12. Keeping up with the research front