Loss functions
Be able to choose a loss for regression and for classification, derive cross-entropy from probability, and explain why MSE is wrong for classes.
Prerequisites
Intuition
A loss measures how wrong the model is — a number the training minimises.
- Regression (predict a number): MSE, the mean squared error. Large errors are punished hard (the square).
- Classification (predict a probability): cross-entropy, −log(the probability the model gave the correct class). If the model gives 0.9 to the right class: loss 0.1. If it gives 0.01: loss 4.6. Being confident and wrong costs enormously — exactly as it should.
Why not MSE on classes? With a sigmoid or softmax the gradient becomes nearly zero when the model is confident and wrong — it does not learn from its worst mistakes. Cross-entropy gives a strong gradient exactly there.
Code
import numpy as np
def mse(y, yhat):
return np.mean((y - yhat) ** 2)
def cross_entropy(p_correct): # p_correct: the model's probability for the right class, per example
return -np.mean(np.log(np.clip(p_correct, 1e-12, 1)))
print(mse(np.array([3.0, 5.0]), np.array([2.5, 6.0]))) # 0.625
print(cross_entropy(np.array([0.9, 0.9, 0.9]))) # 0.105
print(cross_entropy(np.array([0.9, 0.9, 0.01]))) # 1.61 — one confident mistake dominates
In PyTorch: nn.MSELoss(), nn.CrossEntropyLoss() (takes logits and does the softmax + log internally — more numerically stable).
Derivation
Suppose the model gives . Maximum likelihood picks the that maximises , that is, minimises — which is cross-entropy.
For regression under the assumption , : — minimising it gives MSE. So both are the same principle with different noise assumptions.
The gradient for sigmoid + cross-entropy: — linear in the error. For sigmoid + MSE: — the factor as or : the gradient dies exactly when the model is confident and wrong.
Mastery means
- Chooses MSE for regression and cross-entropy for classification, with justification
- Derives cross-entropy from maximum likelihood
- Explains why MSE on probabilities gives weak gradients
Sign in to do the exercises and build your mastery up.
Sources
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Cross-entropy (CC BY-SA 4.0) — CC BY-SA 4.0
Part of the goals (22)
- Language models in practice
- Multimodal systems
- Fine-tune and run your own models
- Build a RAG system you can trust
- Generative models in depth
- Responsible AI in practice
- Build a voice interface
- AI safety in practice
- Fine-tune a model with LoRA
- Evals in practice
- Build an agent you can trust
- Build an NLP system end to end
- AI in production
- An AI service in operation
- Run models more cheaply: quantisation
- Build a memory system for an agent
- Build an AI service that survives production
- Deep reinforcement learning
- Frontier Lab — an independent research project
- Reproduce a paper
- Interpreting a language model
- Statistics for experiments