Skip to content
AI-grafen
EUniversityMathematics· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Information theory: entropy and KL divergence

Be able to compute entropy, cross-entropy and KL divergence and explain why cross-entropy is a natural loss.

Prerequisites

Intuition

Information is surprise. A message that tells you something you already knew carries no information.

Shannon's definition: the surprise of an event with probability pp is −log⁡2p-\log_2 p bits.

ppSurprise
1 (certain)0 bits
0.51 bit
0.252 bits
0.00110 bits

Entropy is the average surprise — how unpredictable a distribution is:

H(p)=−∑ipilog⁡2piH(p) = -\sum_i p_i \log_2 p_i

DistributionEntropyWhy
A coin (0.5, 0.5)1 bitmaximally unpredictable
A biased coin (0.9, 0.1)0.47 bitsfairly predictable
Always the same (1, 0)0 bitsno uncertainty
A die (1/6 × 6)2.58 bitslog₂6

The entropy is also the theoretical lower bound on how much data can be compressed: you cannot pack something into fewer bits than its entropy.

Formal

Three quantities, built on each other:

QuantityFormulaMeans
EntropyH(p)=−∑pilog⁡piH(p) = -\sum p_i \log p_ithe uncertainty in pp
Cross-entropyH(p,q)=−∑pilog⁡qiH(p, q) = -\sum p_i \log q_ithe cost of coding pp with codes built for qq
KL divergenceDKL(p∥q)=∑pilog⁡piqiD_{KL}(p \| q) = \sum p_i \log\frac{p_i}{q_i}the excess cost

The relationship is exact:

H(p,q)=H(p)+DKL(p∥q)H(p, q) = H(p) + D_{KL}(p \| q)

This is the whole explanation of why cross-entropy is ML's standard loss. In classification pp is the true labels — fixed, with an entropy H(p)H(p) you cannot affect. Minimising the cross-entropy is therefore exactly the same thing as minimising the KL divergence between the model and the truth.

And since DKL≥0D_{KL} \geq 0 with equality only when p=qp = q, the minimum is reached precisely when the model has the right distribution. The loss is not chosen out of convenience — it measures the distance to the truth.

Three properties of KL that surprise people:

  1. It is not symmetric. DKL(p∥q)≠DKL(q∥p)D_{KL}(p\|q) \neq D_{KL}(q\|p). That is why it is not a distance in the ordinary sense.
  2. The direction matters. DKL(p∥q)D_{KL}(p\|q) punishes hard if qq gives near zero where pp has mass — the model must not rule out something that actually happens. The reverse direction instead gives the model permission to focus on one peak and ignore the rest. (That is the difference between «mode-covering» and «mode-seeking».)
  3. It becomes infinite if qi=0q_i = 0 where pi>0p_i > 0. That is why a small number is always added in implementations.

Mutual information I(X;Y)=H(X)−H(X∣Y)I(X;Y) = H(X) - H(X\mid Y) measures how much knowledge of YY reduces the uncertainty about XX. It is used to select features and to measure dependencies that a correlation coefficient misses.

Code

import numpy as np

def H(p, base=2):
    p = np.asarray(p, float)
    p = p[p > 0]
    return float(-(p * np.log(p)).sum() / np.log(base))

def cross_entropy(p, q, base=2):
    p, q = np.asarray(p, float), np.asarray(q, float) + 1e-12
    return float(-(p * np.log(q)).sum() / np.log(base))

def kl(p, q, base=2):
    return cross_entropy(p, q, base) - H(p, base)

print(round(H([0.5, 0.5]), 3))          # 1.0    a coin
print(round(H([0.9, 0.1]), 3))          # 0.469  a biased coin
print(round(H([1.0, 0.0]), 3))          # 0.0    no uncertainty
print(round(H([1/6] * 6), 3))           # 2.585  a die

# The relationship H(p,q) = H(p) + KL(p||q)
p, q = [0.7, 0.2, 0.1], [0.5, 0.3, 0.2]
print(round(cross_entropy(p, q), 4), round(H(p) + kl(p, q), 4))   # 1.2796 1.2796

# KL is NOT symmetric
print(round(kl(p, q), 4), round(kl(q, p), 4))                     # 0.1228 0.1328

# Classification: p is one-hot, so H(p) = 0 and cross-entropy = KL
target = [0, 1, 0]
for name, pred in (("confident and right", [0.02, 0.96, 0.02]),
                   ("uncertain", [0.3, 0.4, 0.3]),
                   ("confident and wrong", [0.96, 0.02, 0.02])):
    print(f"  {name:<21} cross-entropy {cross_entropy(target, pred):.3f} bits")
#   confident and right   cross-entropy 0.059 bits
#   uncertain             cross-entropy 1.322 bits
#   confident and wrong   cross-entropy 5.644 bits   ← punished the hardest

# Perplexity: 2^H, «how many alternatives the model is effectively choosing between»
for h in (1.0, 2.0, 3.32):
    print(f"  entropy {h} bits → perplexity {2 ** h:.1f}")

Mastery means

  • Computes entropy and cross-entropy
  • Interprets KL divergence
  • Explains why cross-entropy is the natural loss

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences