Skip to content
AI-grafen
EUniversityStatistics and probability· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

A/B tests and experiment design

Be able to design a controlled experiment, compute the sample size and avoid p-hacking.

Prerequisites

Intuition

An A/B test randomises users to two variants and measures a predetermined metric. The randomisation is what makes the causal claim valid — everything else is on average equal between the groups.

Before the test starts you should have decided:

  1. The primary metric — a single one. (Secondary ones may be reported, but they do not decide.)
  2. The minimum detectable effect (MDE) — how large a difference is worth acting on?
  3. The sample size from the MDE, the baseline, α = 0.05 and a power of 0.8.
  4. The length of the test — at least a whole week (weekday effects).
  5. The stopping rule.

Write it down. Otherwise every result is constructed after the fact.

Formal

The sample size per group for detecting an absolute difference δ\delta in proportions around a baseline pp:

n≈2 (z1−α/2+z1−β)2  p(1−p)δ2n \approx \frac{2\,(z_{1-\alpha/2} + z_{1-\beta})^2\;p(1-p)}{\delta^2}

With α=0.05\alpha = 0.05 (z = 1.96) and a power of 0.8 (z = 0.84) the constant becomes (1.96+0.84)2≈7.85(1.96+0.84)^2 \approx 7.85.

An example: p=0.10p = 0.10, wanting to detect δ=0.01\delta = 0.01 (10 % → 11 %): n≈2⋅7.85⋅0.09/0.0001≈n \approx 2 \cdot 7.85 \cdot 0.09 / 0.0001 \approx 14 100 per group.

Three traps:

  • Peeking: look every day and stop when p < 0.05 → the actual false-alarm risk becomes 20–30 %, not 5 %. The solution: a predetermined length, or sequential tests (always-valid p-values).
  • Multiple metrics: 20 metrics give on average one «significant» one by chance. Correct for it or predetermine one.
  • Interference: users affect each other (social networks, marketplaces) → randomise at the group or market level instead.

Code

from math import ceil

def n_per_group(baseline, mde_absolute, alpha=0.05, power=0.8):
    z = {0.05: 1.96}[alpha] + {0.8: 0.84, 0.9: 1.28}[power]
    return ceil(2 * z**2 * baseline * (1 - baseline) / mde_absolute**2)

for mde in (0.005, 0.01, 0.02):
    print(f"MDE {mde:.1%}: {n_per_group(0.10, mde):>7,} per group")
# MDE 0.5%:  56,376 per group
# MDE 1.0%:  14,094 per group
# MDE 2.0%:   3,524 per group

Halve the MDE → quadruple the n. That is why small improvements require a lot of traffic — and why teams with little traffic should test large changes, not button colours.

Mastery means

  • Designs an A/B test with a predetermined metric
  • Computes the necessary sample size
  • Avoids p-hacking and peeking

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences