Random seeds and the variance between runs
Be able to measure the variance across seeds and report the mean ± the spread.
Prerequisites
- DSamples and uncertaintyrequired
- EReproducibilityrequired
Intuition
Run the same training twice with different random seeds and you get different results. The seed affects the weight initialisation, the data order, the dropout masks and the augmentation.
How much? On small datasets often 1–3 percentage points in the test result. That means a reported improvement of 1 percentage point from a single seed is zero information.
The minimum requirement in an honest report: at least 3 seeds, preferably 5, and the mean ± the standard deviation. Not «the best run».
Code
import numpy as np
def run_with_seeds(train_fn, seeds=(0, 1, 2, 3, 4)):
results = np.array([train_fn(seed=s) for s in seeds])
return {"mean": float(results.mean()), "std": float(results.std(ddof=1)),
"min": float(results.min()), "max": float(results.max()), "all": results.round(4).tolist()}
base = run_with_seeds(lambda seed: train_baseline(seed))
new = run_with_seeds(lambda seed: train_new_method(seed))
print(f"baseline {base['mean']:.3f} ± {base['std']:.3f} new {new['mean']:.3f} ± {new['std']:.3f}")
# baseline 0.842 ± 0.011 new 0.851 ± 0.013
# Is the difference larger than the noise? Paired over the same seeds:
from scipy import stats
d = np.array(new["all"]) - np.array(base["all"])
print(d.mean().round(4), stats.ttest_rel(new["all"], base["all"]).pvalue.round(3))
# 0.009 0.21 → not settled with 5 seeds
How many seeds are needed? Roughly: to detect a difference when the spread is you need about per variant. With σ = 0.012 and δ = 0.01 that is ~23 seeds. If you do not have that budget, the honest conclusion is «the difference is smaller than we can measure» — which is itself a result worth reporting.
Mastery means
- Measures the variance across seeds
- Reports the mean ± the spread
- Decides how many seeds are needed
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Show Your Work: Improved Reporting of Experimental Results — arXiv (open access; licence per article)
- arXiv — Deep Reinforcement Learning that Matters — arXiv (open access; licence per article)