EUniversityStatistics and probability· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN
Hypothesis testing and p-values
Be able to set up a null and an alternative hypothesis, interpret p-values and understand their limitations.
Prerequisites
- EConfidence intervalsrequired
Intuition
The null hypothesis H0: «no difference». The p-value = the probability of seeing a difference this large (or larger) if H0 were true. A small p (< 0.05) → the data is unlikely under H0 → we reject H0.
What p is not: the probability that H0 is true; a measure of how large the difference is; a proof.
In ML:
- Two models on the same test cases → McNemar (classification) or a paired t-test/permutation test on the per-case differences.
- Many comparisons (10 prompts × 5 metrics) → some become «significant» by chance. Bonferroni (α/m) or report honestly how many tests were run.
- Report the effect size with an interval — «+2.1 pp [0.4, 3.8]» says more than «p = 0.03».
With n = 100 000 everything becomes significant; with n = 30 nothing does. p measures the evidence and the sample size.
Code
import numpy as np
from scipy import stats
# a permutation test on paired differences (it works for any metric at all)
def perm_test(diff, n=20000, seed=0):
rng = np.random.default_rng(seed); diff = np.asarray(diff); obs = diff.mean()
flips = rng.choice([-1, 1], size=(n, len(diff)))
null = (flips * diff).mean(axis=1)
return obs, np.mean(np.abs(null) >= abs(obs))
right_a = np.array([1,1,0,1,1,1,0,1,1,1,0,1,1,1,1,0,1,1,1,1,1,0,1,1,1,1,1,0,1,1])
right_b = np.array([1,0,0,1,1,1,0,1,0,1,0,1,1,1,1,0,1,0,1,1,1,0,1,0,1,1,0,0,1,1])
print(perm_test(right_a - right_b)) # (0.167, ~0.03)
# McNemar: only the discordant cases count
b = int(((right_a == 1) & (right_b == 0)).sum()); c = int(((right_a == 0) & (right_b == 1)).sum())
print(b, c, stats.binomtest(b, b + c, 0.5).pvalue) # 5 0 → p = 0.0625 (exact)
# Bonferroni: 12 comparisons → a threshold of 0.05/12
Mastery means
- Sets up H0/H1 and chooses a test (t-test, McNemar, permutation)
- Interprets a p-value correctly and knows its limitations
- Corrects for multiple comparisons and reports the effect size
Sign in to do the exercises and build your mastery up.
Sources
- Swedish Wikipedia — p-value (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — McNemar's test (CC BY-SA 4.0) — CC BY-SA 4.0
- ASA statement on p-values (2016) — open PDF