Information theory: entropy and KL divergence
Be able to compute entropy, cross-entropy and KL divergence and explain why cross-entropy is a natural loss.
Prerequisites
Intuition
Information is surprise. A message that tells you something you already knew carries no information.
Shannon's definition: the surprise of an event with probability is bits.
| Surprise | |
|---|---|
| 1 (certain) | 0 bits |
| 0.5 | 1 bit |
| 0.25 | 2 bits |
| 0.001 | 10 bits |
Entropy is the average surprise — how unpredictable a distribution is:
| Distribution | Entropy | Why |
|---|---|---|
| A coin (0.5, 0.5) | 1 bit | maximally unpredictable |
| A biased coin (0.9, 0.1) | 0.47 bits | fairly predictable |
| Always the same (1, 0) | 0 bits | no uncertainty |
| A die (1/6 × 6) | 2.58 bits | log₂6 |
The entropy is also the theoretical lower bound on how much data can be compressed: you cannot pack something into fewer bits than its entropy.
Formal
Three quantities, built on each other:
| Quantity | Formula | Means |
|---|---|---|
| Entropy | the uncertainty in | |
| Cross-entropy | the cost of coding with codes built for | |
| KL divergence | the excess cost |
The relationship is exact:
This is the whole explanation of why cross-entropy is ML's standard loss. In classification is the true labels — fixed, with an entropy you cannot affect. Minimising the cross-entropy is therefore exactly the same thing as minimising the KL divergence between the model and the truth.
And since with equality only when , the minimum is reached precisely when the model has the right distribution. The loss is not chosen out of convenience — it measures the distance to the truth.
Three properties of KL that surprise people:
- It is not symmetric. . That is why it is not a distance in the ordinary sense.
- The direction matters. punishes hard if gives near zero where has mass — the model must not rule out something that actually happens. The reverse direction instead gives the model permission to focus on one peak and ignore the rest. (That is the difference between «mode-covering» and «mode-seeking».)
- It becomes infinite if where . That is why a small number is always added in implementations.
Mutual information measures how much knowledge of reduces the uncertainty about . It is used to select features and to measure dependencies that a correlation coefficient misses.
Code
import numpy as np
def H(p, base=2):
p = np.asarray(p, float)
p = p[p > 0]
return float(-(p * np.log(p)).sum() / np.log(base))
def cross_entropy(p, q, base=2):
p, q = np.asarray(p, float), np.asarray(q, float) + 1e-12
return float(-(p * np.log(q)).sum() / np.log(base))
def kl(p, q, base=2):
return cross_entropy(p, q, base) - H(p, base)
print(round(H([0.5, 0.5]), 3)) # 1.0 a coin
print(round(H([0.9, 0.1]), 3)) # 0.469 a biased coin
print(round(H([1.0, 0.0]), 3)) # 0.0 no uncertainty
print(round(H([1/6] * 6), 3)) # 2.585 a die
# The relationship H(p,q) = H(p) + KL(p||q)
p, q = [0.7, 0.2, 0.1], [0.5, 0.3, 0.2]
print(round(cross_entropy(p, q), 4), round(H(p) + kl(p, q), 4)) # 1.2796 1.2796
# KL is NOT symmetric
print(round(kl(p, q), 4), round(kl(q, p), 4)) # 0.1228 0.1328
# Classification: p is one-hot, so H(p) = 0 and cross-entropy = KL
target = [0, 1, 0]
for name, pred in (("confident and right", [0.02, 0.96, 0.02]),
("uncertain", [0.3, 0.4, 0.3]),
("confident and wrong", [0.96, 0.02, 0.02])):
print(f" {name:<21} cross-entropy {cross_entropy(target, pred):.3f} bits")
# confident and right cross-entropy 0.059 bits
# uncertain cross-entropy 1.322 bits
# confident and wrong cross-entropy 5.644 bits ← punished the hardest
# Perplexity: 2^H, «how many alternatives the model is effectively choosing between»
for h in (1.0, 2.0, 3.32):
print(f" entropy {h} bits → perplexity {2 ** h:.1f}")
Mastery means
- Computes entropy and cross-entropy
- Interprets KL divergence
- Explains why cross-entropy is the natural loss
Sign in to do the exercises and build your mastery up.
Sources
- Mathematics for Machine Learning (Deisenroth m.fl.) — free to read online (authors' edition)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Entropy (information theory) (CC BY-SA 4.0) — CC BY-SA 4.0