Reproducing a paper
Be able to reproduce a published result from the article to running code, compare with the reported figures and explain the deviations.
Prerequisites
- EReproducibilityrequired
- FReading and analysing research papersrequired
- GAblation studiesrequired
Intuition
Reproducing a paper is the best way to understand it — and to discover that some of what it says does not hold, or is not said at all.
The workflow:
- Pick a result: one main claim, one table, one curve. Not the whole paper.
- Scale it down: a smaller model or dataset so that a run takes minutes. The effect should still be visible (otherwise: pick a different result).
- List everything missing from the paper: the lr schedule, the seeds, the preprocessing, «details in the appendix» that are not there. Guess, and document the guess.
- Build the baseline first and verify that it roughly matches the paper's baseline. If it does not — stop and understand why.
- Run the method with ≥ 3 seeds. Compare the direction and the size of the effect, not the exact decimals.
- Explain the deviations: the scale, missing details, a bug on your side, a bug on theirs.
- Report: what was reproduced, what was not, and what was needed beyond the paper.
A failed reproduction with clear documentation is a valuable result.
Research
The state of the field: the ML Reproducibility Challenge (since 2018) shows that a substantial share of published results cannot be recreated even with the authors' code — the common causes being underreported hyperparameters, unspecified preprocessing, seed sensitivity and a weak baseline. Henderson et al. (2018) showed that the same RL algorithm with different codebases or seeds gives large differences. The field's response: reproducibility checklists (NeurIPS), code requirements, «Papers with Code», and preregistration.
For your reproduction: use the Papers with Code link if there is one, but write your own implementation of the core — that is where the understanding lies. Then compare against their code: the difference is often exactly where the paper's text is incomplete.
Project 6 does this for Srivastava et al. (2014): an MLP on digits with a small training set, p ∈ {0, 0.2, 0.5}, three seeds, and a report with a hypothesis, a method and a table.
Mastery means
- Goes from the article to running code that recreates the main result at a small scale
- Compares against the reported figures and explains the deviations
- Documents the reproduction so that others can repeat it
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR) — arXiv (open access; licence per article)
- arXiv — Deep Reinforcement Learning that Matters — arXiv (open access; licence per article)
- ML Reproducibility Challenge — free to read