Confidence intervals
Be able to compute and interpret confidence intervals for means and proportions.
Prerequisites
Intuition
A 95 % confidence interval is an interval computed so that, if you repeated the whole sampling procedure many times, 95 % of the intervals would cover the true value. It is not «a 95 % probability that the truth lies here» — the truth is fixed, the interval is what is random.
For a mean: x̄ ± 1.96·s/√n (large n). For a proportion: p̂ ± 1.96·√(p̂(1−p̂)/n). At a small n or proportions near 0/1: the Wilson interval.
The use in ML: report test metrics with an interval; two intervals that overlap a lot → the difference is not shown. The bootstrap gives an interval for any metric at all (F1, BLEU, an agent's solve rate) without a formula.
Code
import numpy as np
from scipy import stats
# a mean, the t distribution (a small n)
errors = np.array([3.1, 2.4, 4.0, 3.3, 2.9, 3.8, 3.5, 2.7]) # the MAE per run
m, se = errors.mean(), errors.std(ddof=1) / np.sqrt(len(errors))
lo, hi = stats.t.interval(0.95, len(errors) - 1, loc=m, scale=se)
print(round(m, 2), round(lo, 2), round(hi, 2)) # 3.21 2.75 3.67
# a proportion: Wilson
def wilson(k, n, z=1.96):
p = k / n; d = 1 + z**2 / n
c = (p + z**2 / (2 * n)) / d; h = z * np.sqrt(p * (1 - p) / n + z**2 / (4 * n**2)) / d
return c - h, c + h
print(wilson(87, 100)) # (0.79, 0.92)
# the bootstrap for F1
from sklearn.metrics import f1_score
rng = np.random.default_rng(0); idx = np.arange(len(y_true))
f1s = [f1_score(y_true[i], y_pred[i], average="macro") for i in (rng.choice(idx, len(idx)) for _ in range(2000))]
print(np.percentile(f1s, [2.5, 97.5]))
Mastery means
- Computes a confidence interval for a mean and for a proportion
- Interprets the interval correctly (frequentist)
- Uses the bootstrap for arbitrary metrics
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Konfidensintervall (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Binomial proportion confidence interval (CC BY-SA 4.0) — CC BY-SA 4.0