Maximum likelihood
Be able to derive ML estimates and explain that minimising cross-entropy is maximising likelihood.
Prerequisites
Intuition
Maximum likelihood answers the question: which parameter values make the data I actually observed the most likely?
You flip a coin 10 times and get 7 heads. Which is the most likely?
| The probability of exactly 7 out of 10 | |
|---|---|
| 0.3 | 0.009 |
| 0.5 | 0.117 |
| 0.7 | 0.267 |
| 0.9 | 0.057 |
The answer is 0.7 — the proportion in the data. It feels obvious, but it follows from the principle rather than being assumed.
And the principle is general. Nearly every loss function in machine learning is a maximum likelihood estimate under some assumption about the noise. It is not a collection of tricks — it is one principle with different special cases.
Derivation
The likelihood is the probability of the data, seen as a function of the parameters:
We maximise the log-likelihood instead, for two reasons: the product becomes a sum (differentiable term by term), and the underflow disappears. The logarithm is increasing, so the maximum is in the same place.
Example 1 — the coin. With heads out of :
Example 2 — the mean of a normal distribution.
Maximising this is minimising the sum of squared deviations. So:
Least squares is maximum likelihood under the assumption that the noise is normally distributed.
Differentiating gives — the sample mean.
Example 3 — classification. With and the model :
The negative of this is binary cross-entropy. So:
Minimising cross-entropy is maximising likelihood under a Bernoulli model.
The table that ties it all together:
| The assumption about the noise | ML gives the loss |
|---|---|
| Normally distributed | MSE |
| Laplace distributed | MAE (absolute error) |
| Bernoulli | binary cross-entropy |
| Categorical | cross-entropy |
| Poisson | Poisson deviance |
And regularisation has the same origin. Add a prior and maximise the posterior instead (MAP) and a Gaussian prior becomes exactly L2 regularisation, and a Laplace prior exactly L1. They are therefore not bolted-on penalties but the consequence of an assumed distribution on the parameters.
ML's properties: consistent (it converges towards the truth with more data) and asymptotically efficient (no estimator is better in the limit). But not always unbiased — the ML estimate of a normal distribution's variance divides by , not , and therefore systematically underestimates.
Code
import numpy as np
from scipy.optimize import minimize_scalar
# 1. The coin: the log-likelihood and its maximum
k, n = 7, 10
def neg_ll(p):
if not 0 < p < 1:
return np.inf
return -(k * np.log(p) + (n - k) * np.log(1 - p))
for p in (0.3, 0.5, 0.7, 0.9):
print(f" p={p} L={np.exp(-neg_ll(p)):.4f}")
print("the ML estimate:", round(minimize_scalar(neg_ll, bounds=(1e-6, 1-1e-6),
method="bounded").x, 4)) # 0.7
# 2. The normal distribution: ML gives the mean, and the variance with n (not n-1)
rng = np.random.default_rng(0)
x = rng.normal(loc=5.0, scale=2.0, size=50)
print(round(x.mean(), 4), round(float(x.var(ddof=0)), 4), round(float(x.var(ddof=1)), 4))
# the ML mean the ML variance (/n) the unbiased one (/n-1)
# The ML variance systematically underestimates — show it by repetition
bias = [rng.normal(5, 2, 20).var(ddof=0) for _ in range(20000)]
print(round(float(np.mean(bias)), 3), "against the true value 4.0") # ~3.80
# 3. Least squares IS maximum likelihood under normal noise
xs = np.array([1.0, 2.0, 3.0, 4.0, 5.0])
ys = np.array([2.1, 3.9, 6.2, 7.8, 10.1])
def neg_ll_linear(params, sigma=1.0):
k_, m_ = params
res = ys - (k_ * xs + m_)
return 0.5 * np.sum(res**2) / sigma**2 + len(xs) * np.log(sigma)
from scipy.optimize import minimize
ml = minimize(neg_ll_linear, [0.0, 0.0]).x
ls = np.polyfit(xs, ys, 1)
print(np.round(ml, 4), np.round(ls, 4)) # [1.99 0.03] [1.99 0.03] — identical
# 4. Cross-entropy IS the negative log-likelihood
y = np.array([1, 0, 1, 1])
phat = np.array([0.9, 0.2, 0.8, 0.6])
bce = -np.mean(y * np.log(phat) + (1 - y) * np.log(1 - phat))
ll = np.sum(y * np.log(phat) + (1 - y) * np.log(1 - phat))
print(round(float(bce), 4), round(float(-ll / len(y)), 4)) # the same number
Mastery means
- Derives an ML estimate
- Uses the log-likelihood and knows why
- Connects ML to cross-entropy and MSE
Sign in to do the exercises and build your mastery up.
Sources
- Mathematics for Machine Learning (Deisenroth m.fl.) — free to read online (authors' edition)
- scikit-learn User Guide (BSD-3) — BSD-3-Clause
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0