A/B tests and experiment design
Be able to design a controlled experiment, compute the sample size and avoid p-hacking.
Prerequisites
- EHypothesis testing and p-valuesrequired
Intuition
An A/B test randomises users to two variants and measures a predetermined metric. The randomisation is what makes the causal claim valid — everything else is on average equal between the groups.
Before the test starts you should have decided:
- The primary metric — a single one. (Secondary ones may be reported, but they do not decide.)
- The minimum detectable effect (MDE) — how large a difference is worth acting on?
- The sample size from the MDE, the baseline, α = 0.05 and a power of 0.8.
- The length of the test — at least a whole week (weekday effects).
- The stopping rule.
Write it down. Otherwise every result is constructed after the fact.
Formal
The sample size per group for detecting an absolute difference in proportions around a baseline :
With (z = 1.96) and a power of 0.8 (z = 0.84) the constant becomes .
An example: , wanting to detect (10 % → 11 %): 14 100 per group.
Three traps:
- Peeking: look every day and stop when p < 0.05 → the actual false-alarm risk becomes 20–30 %, not 5 %. The solution: a predetermined length, or sequential tests (always-valid p-values).
- Multiple metrics: 20 metrics give on average one «significant» one by chance. Correct for it or predetermine one.
- Interference: users affect each other (social networks, marketplaces) → randomise at the group or market level instead.
Code
from math import ceil
def n_per_group(baseline, mde_absolute, alpha=0.05, power=0.8):
z = {0.05: 1.96}[alpha] + {0.8: 0.84, 0.9: 1.28}[power]
return ceil(2 * z**2 * baseline * (1 - baseline) / mde_absolute**2)
for mde in (0.005, 0.01, 0.02):
print(f"MDE {mde:.1%}: {n_per_group(0.10, mde):>7,} per group")
# MDE 0.5%: 56,376 per group
# MDE 1.0%: 14,094 per group
# MDE 2.0%: 3,524 per group
Halve the MDE → quadruple the n. That is why small improvements require a lot of traffic — and why teams with little traffic should test large changes, not button colours.
Mastery means
- Designs an A/B test with a predetermined metric
- Computes the necessary sample size
- Avoids p-hacking and peeking
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — A/B testing (CC BY-SA 4.0) — CC BY-SA 4.0
- Kohavi m.fl. — Trustworthy Online Controlled Experiments (referens) — free to read