Skip to content
AI-grafen
EUniversityStatistics and probability· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Maximum likelihood

Be able to derive ML estimates and explain that minimising cross-entropy is maximising likelihood.

Prerequisites

Intuition

Maximum likelihood answers the question: which parameter values make the data I actually observed the most likely?

You flip a coin 10 times and get 7 heads. Which pp is the most likely?

ppThe probability of exactly 7 out of 10
0.30.009
0.50.117
0.70.267
0.90.057

The answer is 0.7 — the proportion in the data. It feels obvious, but it follows from the principle rather than being assumed.

And the principle is general. Nearly every loss function in machine learning is a maximum likelihood estimate under some assumption about the noise. It is not a collection of tricks — it is one principle with different special cases.

Derivation

The likelihood is the probability of the data, seen as a function of the parameters:

L(θ)=∏i=1np(xi∣θ)L(\theta) = \prod_{i=1}^{n} p(x_i \mid \theta)

We maximise the log-likelihood instead, for two reasons: the product becomes a sum (differentiable term by term), and the underflow disappears. The logarithm is increasing, so the maximum is in the same place.

Example 1 — the coin. With kk heads out of nn:

ℓ(p)=klog⁡p+(n−k)log⁡(1−p)\ell(p) = k\log p + (n-k)\log(1-p) dℓdp=kp−n−k1−p=0  ⟹  p^=kn\frac{d\ell}{dp} = \frac{k}{p} - \frac{n-k}{1-p} = 0 \;\Longrightarrow\; \hat{p} = \frac{k}{n}

Example 2 — the mean of a normal distribution.

ℓ(μ)=−12σ2∑i(xi−μ)2+a constant\ell(\mu) = -\frac{1}{2\sigma^2}\sum_i (x_i - \mu)^2 + \text{a constant}

Maximising this is minimising the sum of squared deviations. So:

Least squares is maximum likelihood under the assumption that the noise is normally distributed.

Differentiating gives μ^=xˉ\hat{\mu} = \bar{x} — the sample mean.

Example 3 — classification. With yi∈{0,1}y_i \in \{0,1\} and the model p^i\hat{p}_i:

ℓ=∑i[yilog⁡p^i+(1−yi)log⁡(1−p^i)]\ell = \sum_i \left[y_i\log\hat{p}_i + (1-y_i)\log(1-\hat{p}_i)\right]

The negative of this is binary cross-entropy. So:

Minimising cross-entropy is maximising likelihood under a Bernoulli model.

The table that ties it all together:

The assumption about the noiseML gives the loss
Normally distributedMSE
Laplace distributedMAE (absolute error)
Bernoullibinary cross-entropy
Categoricalcross-entropy
PoissonPoisson deviance

And regularisation has the same origin. Add a prior and maximise the posterior instead (MAP) and a Gaussian prior becomes exactly L2 regularisation, and a Laplace prior exactly L1. They are therefore not bolted-on penalties but the consequence of an assumed distribution on the parameters.

ML's properties: consistent (it converges towards the truth with more data) and asymptotically efficient (no estimator is better in the limit). But not always unbiased — the ML estimate of a normal distribution's variance divides by nn, not n−1n-1, and therefore systematically underestimates.

Code

import numpy as np
from scipy.optimize import minimize_scalar

# 1. The coin: the log-likelihood and its maximum
k, n = 7, 10

def neg_ll(p):
    if not 0 < p < 1:
        return np.inf
    return -(k * np.log(p) + (n - k) * np.log(1 - p))

for p in (0.3, 0.5, 0.7, 0.9):
    print(f"  p={p}  L={np.exp(-neg_ll(p)):.4f}")
print("the ML estimate:", round(minimize_scalar(neg_ll, bounds=(1e-6, 1-1e-6),
                                                method="bounded").x, 4))   # 0.7

# 2. The normal distribution: ML gives the mean, and the variance with n (not n-1)
rng = np.random.default_rng(0)
x = rng.normal(loc=5.0, scale=2.0, size=50)
print(round(x.mean(), 4), round(float(x.var(ddof=0)), 4), round(float(x.var(ddof=1)), 4))
#            the ML mean   the ML variance (/n)   the unbiased one (/n-1)

# The ML variance systematically underestimates — show it by repetition
bias = [rng.normal(5, 2, 20).var(ddof=0) for _ in range(20000)]
print(round(float(np.mean(bias)), 3), "against the true value 4.0")     # ~3.80

# 3. Least squares IS maximum likelihood under normal noise
xs = np.array([1.0, 2.0, 3.0, 4.0, 5.0])
ys = np.array([2.1, 3.9, 6.2, 7.8, 10.1])

def neg_ll_linear(params, sigma=1.0):
    k_, m_ = params
    res = ys - (k_ * xs + m_)
    return 0.5 * np.sum(res**2) / sigma**2 + len(xs) * np.log(sigma)

from scipy.optimize import minimize
ml = minimize(neg_ll_linear, [0.0, 0.0]).x
ls = np.polyfit(xs, ys, 1)
print(np.round(ml, 4), np.round(ls, 4))    # [1.99 0.03] [1.99 0.03] — identical

# 4. Cross-entropy IS the negative log-likelihood
y = np.array([1, 0, 1, 1])
phat = np.array([0.9, 0.2, 0.8, 0.6])
bce = -np.mean(y * np.log(phat) + (1 - y) * np.log(1 - phat))
ll = np.sum(y * np.log(phat) + (1 - y) * np.log(1 - phat))
print(round(float(bce), 4), round(float(-ll / len(y)), 4))     # the same number

Mastery means

  • Derives an ML estimate
  • Uses the log-likelihood and knows why
  • Connects ML to cross-entropy and MSE

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences