Skip to content
AI-grafen
GFrontier LabScientific method· about 240 min· evolving, reviewed regularly· verified 2026-09-20· EN

Reproducing a paper

Be able to reproduce a published result from the article to running code, compare with the reported figures and explain the deviations.

Prerequisites

Intuition

Reproducing a paper is the best way to understand it — and to discover that some of what it says does not hold, or is not said at all.

The workflow:

  1. Pick a result: one main claim, one table, one curve. Not the whole paper.
  2. Scale it down: a smaller model or dataset so that a run takes minutes. The effect should still be visible (otherwise: pick a different result).
  3. List everything missing from the paper: the lr schedule, the seeds, the preprocessing, «details in the appendix» that are not there. Guess, and document the guess.
  4. Build the baseline first and verify that it roughly matches the paper's baseline. If it does not — stop and understand why.
  5. Run the method with ≥ 3 seeds. Compare the direction and the size of the effect, not the exact decimals.
  6. Explain the deviations: the scale, missing details, a bug on your side, a bug on theirs.
  7. Report: what was reproduced, what was not, and what was needed beyond the paper.

A failed reproduction with clear documentation is a valuable result.

Research

The state of the field: the ML Reproducibility Challenge (since 2018) shows that a substantial share of published results cannot be recreated even with the authors' code — the common causes being underreported hyperparameters, unspecified preprocessing, seed sensitivity and a weak baseline. Henderson et al. (2018) showed that the same RL algorithm with different codebases or seeds gives large differences. The field's response: reproducibility checklists (NeurIPS), code requirements, «Papers with Code», and preregistration.

For your reproduction: use the Papers with Code link if there is one, but write your own implementation of the core — that is where the understanding lies. Then compare against their code: the difference is often exactly where the paper's text is incomplete.

Project 6 does this for Srivastava et al. (2014): an MLP on digits with a small training set, p ∈ {0, 0.2, 0.5}, three seeds, and a report with a hypothesis, a method and a table.

Mastery means

  • Goes from the article to running code that recreates the main result at a small scale
  • Compares against the reported figures and explains the deviations
  • Documents the reproduction so that others can repeat it

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences