Skip to content
AI-grafen
DAI developerDeep learning· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Loss functions

Be able to choose a loss for regression and for classification, derive cross-entropy from probability, and explain why MSE is wrong for classes.

Prerequisites

Intuition

A loss measures how wrong the model is — a number the training minimises.

  • Regression (predict a number): MSE, the mean squared error. Large errors are punished hard (the square).
  • Classification (predict a probability): cross-entropy, −log(the probability the model gave the correct class). If the model gives 0.9 to the right class: loss 0.1. If it gives 0.01: loss 4.6. Being confident and wrong costs enormously — exactly as it should.

Why not MSE on classes? With a sigmoid or softmax the gradient becomes nearly zero when the model is confident and wrong — it does not learn from its worst mistakes. Cross-entropy gives a strong gradient exactly there.

Code

import numpy as np

def mse(y, yhat):
    return np.mean((y - yhat) ** 2)

def cross_entropy(p_correct):       # p_correct: the model's probability for the right class, per example
    return -np.mean(np.log(np.clip(p_correct, 1e-12, 1)))

print(mse(np.array([3.0, 5.0]), np.array([2.5, 6.0])))      # 0.625
print(cross_entropy(np.array([0.9, 0.9, 0.9])))            # 0.105
print(cross_entropy(np.array([0.9, 0.9, 0.01])))           # 1.61 — one confident mistake dominates

In PyTorch: nn.MSELoss(), nn.CrossEntropyLoss() (takes logits and does the softmax + log internally — more numerically stable).

Derivation

Suppose the model gives pθ(y∣x)p_\theta(y\mid x). Maximum likelihood picks the θ\theta that maximises ∏ipθ(yi∣xi)\prod_i p_\theta(y_i\mid x_i), that is, minimises −∑ilog⁡pθ(yi∣xi)-\sum_i \log p_\theta(y_i\mid x_i) — which is cross-entropy.

For regression under the assumption y=fθ(x)+εy = f_\theta(x) + \varepsilon, ε∼N(0,σ2)\varepsilon\sim\mathcal N(0,\sigma^2): −log⁡p=(y−fθ(x))22σ2+const-\log p = \frac{(y - f_\theta(x))^2}{2\sigma^2} + \text{const} — minimising it gives MSE. So both are the same principle with different noise assumptions.

The gradient for sigmoid + cross-entropy: ∂L/∂z=p^−y\partial L/\partial z = \hat p - y — linear in the error. For sigmoid + MSE: (p^−y) p^(1−p^)(\hat p - y)\,\hat p(1-\hat p) — the factor p^(1−p^)→0\hat p(1-\hat p)\to 0 as p^→0\hat p\to 0 or 11: the gradient dies exactly when the model is confident and wrong.

Mastery means

  • Chooses MSE for regression and cross-entropy for classification, with justification
  • Derives cross-entropy from maximum likelihood
  • Explains why MSE on probabilities gives weak gradients

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences