Integrals in probability: the expectation
Be able to compute the expectation and the variance for continuous distributions.
Prerequisites
Intuition
For a discrete distribution the expectation is a sum: E[X] = Σ x·P(x). For a continuous distribution the sum becomes an integral:
The same idea: every value is weighted by its probability density.
The variance measures the spread:
The second form is nearly always easier to compute with.
Formal
Uniform on [a, b]: , , .
Exponential with parameter λ: for , , .
Normal: , .
Linearity — the most used property: holds always, even when X and Y are dependent. The variance, on the other hand, is not linear: The last term disappears only if X and Y are uncorrelated.
The connection to ML: every loss function is an expectation over the data distribution, . We cannot compute the integral — we estimate it with a sample, which is exactly what a minibatch is. That of the minibatch gradient is the true gradient (unbiased) is the whole basis for SGD working.
Code
import numpy as np
from scipy import integrate
# Exponential with lambda = 2: E[X] = 0.5, Var = 0.25
p = lambda x: 2 * np.exp(-2 * x)
E = integrate.quad(lambda x: x * p(x), 0, np.inf)[0]
E2 = integrate.quad(lambda x: x**2 * p(x), 0, np.inf)[0]
print(round(E, 4), round(E2 - E**2, 4)) # 0.5 0.25
# Monte Carlo — the same answer without solving the integral
rng = np.random.default_rng(0)
x = rng.exponential(scale=0.5, size=200_000)
print(round(x.mean(), 4), round(x.var(), 4)) # 0.4997 0.2494
# The minibatch gradient is an unbiased estimate of the true one
def true_gradient(data, theta):
return np.mean([grad(d, theta) for d in data], axis=0)
def minibatch_gradient(data, theta, bs, rng):
return np.mean([grad(d, theta) for d in rng.choice(data, bs, replace=False)], axis=0)
# E[minibatch_gradient] == true_gradient → that is why SGD converges
Mastery means
- Computes the expectation and the variance for continuous distributions
- Connects the integral to a Monte Carlo estimate
- Uses linearity
Sign in to do the exercises and build your mastery up.
Sources
- Swedish Wikipedia — Expected value (CC BY-SA 4.0) — CC BY-SA 4.0
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0