Statistics — mean, median and spread
Be able to calculate the mean, the median and the standard deviation of a small dataset, choose the right measure for the situation, and explain why spread matters as much as the middle.
Practise in Mattegrafen ↗ · Sannolikhet och statistikPrerequisites
Intuition
Five colleagues' monthly commuting costs: 200, 250, 250, 300, 3000 (one of them flies). The mean is 800 kr — but nobody gets anywhere near that. The median (the middle value) is 250 and describes the group better. An extreme value drags the mean but not the median.
Spread says how spread out the values are. Two classes can have the same average grade, but one has everybody in the middle and the other has half at the top and half at the bottom.
Formal
For the values x₁ … xₙ:
- Mean: x̄ = (x₁ + … + xₙ) / n
- Variance: s² = Σ(xᵢ − x̄)² / n (population variance)
- Standard deviation: s = √s²
Example [2, 4, 4, 4, 5, 5, 7, 9]: x̄ = 5, the squared deviations sum to 32, s² = 4, s = 2.
Why in AI? A model's errors are a dataset. The mean error says how good it is on average; the spread says whether it is evenly good or occasionally catastrophic. And normalising the data (subtract the mean, divide by s) is often the very first step before training.
Mastery means
- Calculates the mean and the median correctly
- Explains when the median is better than the mean
Sign in to do the exercises and build your mastery up.
Sources
Leads to
- CLinear regression: fitting a straight line to data
- CProbability — the basics
- DBetter or just different?
- DClustering: k-means and hierarchical
- DCorrelation and causation
- DLinear regression with several features
- DProbability distributions
- EModel evaluation
- EPCA — principal component analysis
- EScientific method in AI
Part of the goals (34)
- The mathematics behind the models
- Build an NLP system end to end
- Statistics for experiments
- Classical machine learning in practice
- Fine-tune a model with LoRA
- Build a RAG system you can trust
- AI in production
- Generative models in depth
- Classical ML for real
- Training neural networks for real
- Frontier Lab — an independent research project
- Evals in practice
- Language models in practice
- Understand how generative AI works
- Train an agent with reward
- Fine-tune and run your own models
- Build a memory system for an agent
- Build an agent you can trust
- AI safety in practice
- Responsible AI in practice
- Build an AI service that survives production
- An AI service in operation
- Interpreting a language model
- Reproduce a paper
- Build a voice interface
- Train your first neural network
- Deep reinforcement learning
- Seeing and hearing with AI
- Data: collect, clean, document
- Image classification with convolutional networks
- Run models more cheaply: quantisation
- Multimodal systems
- Build a transformer from scratch
- AI, ethics and society