Skip to content
AI-grafen
EUniversityStatistics and probability· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Hypothesis testing and p-values

Be able to set up a null and an alternative hypothesis, interpret p-values and understand their limitations.

Prerequisites

Intuition

The null hypothesis H0: «no difference». The p-value = the probability of seeing a difference this large (or larger) if H0 were true. A small p (< 0.05) → the data is unlikely under H0 → we reject H0.

What p is not: the probability that H0 is true; a measure of how large the difference is; a proof.

In ML:

  • Two models on the same test cases → McNemar (classification) or a paired t-test/permutation test on the per-case differences.
  • Many comparisons (10 prompts × 5 metrics) → some become «significant» by chance. Bonferroni (α/m) or report honestly how many tests were run.
  • Report the effect size with an interval — «+2.1 pp [0.4, 3.8]» says more than «p = 0.03».

With n = 100 000 everything becomes significant; with n = 30 nothing does. p measures the evidence and the sample size.

Code

import numpy as np
from scipy import stats

# a permutation test on paired differences (it works for any metric at all)
def perm_test(diff, n=20000, seed=0):
    rng = np.random.default_rng(seed); diff = np.asarray(diff); obs = diff.mean()
    flips = rng.choice([-1, 1], size=(n, len(diff)))
    null = (flips * diff).mean(axis=1)
    return obs, np.mean(np.abs(null) >= abs(obs))

right_a = np.array([1,1,0,1,1,1,0,1,1,1,0,1,1,1,1,0,1,1,1,1,1,0,1,1,1,1,1,0,1,1])
right_b = np.array([1,0,0,1,1,1,0,1,0,1,0,1,1,1,1,0,1,0,1,1,1,0,1,0,1,1,0,0,1,1])
print(perm_test(right_a - right_b))                     # (0.167, ~0.03)

# McNemar: only the discordant cases count
b = int(((right_a == 1) & (right_b == 0)).sum()); c = int(((right_a == 0) & (right_b == 1)).sum())
print(b, c, stats.binomtest(b, b + c, 0.5).pvalue)      # 5 0 → p = 0.0625 (exact)

# Bonferroni: 12 comparisons → a threshold of 0.05/12

Mastery means

  • Sets up H0/H1 and chooses a test (t-test, McNemar, permutation)
  • Interprets a p-value correctly and knows its limitations
  • Corrects for multiple comparisons and reports the effect size

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences