Skip to content
AI-grafen
GFrontier LabProject· about 240 min· server sandbox

Project 6: Reproduce a paper — Dropout (Srivastava et al. 2014)

Formulate the paper's claim as a hypothesis, reproduce it at small scale (dropout reduces the gap between training and test error), run an ablation over p ∈ {0, 0.2, 0.5} and three seeds, report mean ± std and write a reproducible report as an artifact.

Theory

Srivastava et al. (2014) show that dropout reduces overfitting by randomly switching off neurons during training. Hypothesis: with dropout p=0.5 the test−train gap is smaller than without, and the test error lower, on a small MLP trained on little data. Reproduction requires: fixed seeds, the same data, the same budget, several runs.

Sub-tasks

  1. model with dropout — make_model(p) → MLP 64→256→256→10 with nn.Dropout(p) after every ReLU.
  2. one experiment — run_once(p, seed, n_train, epochs) → {train_acc, test_acc, gap} on digits with a small training set (n_train=300) to provoke overfitting.
  3. ablation — ablation(ps, seeds) → per p: mean and std of gap and test_acc.
  4. report — write_report(results, path) writes out/report.md with a table, hypothesis, method, seeds and conclusion.

Passes when: gap_reduction >= 0.15

The starter code

runs in an isolated sandbox on the server
import os
import statistics
import torch
import torch.nn as nn
from sklearn.datasets import load_digits


def data(seed, n_train=300):
    d = load_digits()
    X = torch.tensor(d.data, dtype=torch.float32) / 16; y = torch.tensor(d.target)
    g = torch.Generator().manual_seed(seed); idx = torch.randperm(len(X), generator=g)
    return X[idx[:n_train]], y[idx[:n_train]], X[idx[n_train:]], y[idx[n_train:]]


def make_model(p):
    # TODO
    ...


def accuracy(model, X, y):
    model.eval()
    with torch.no_grad():
        return float((model(X).argmax(1) == y).float().mean())


def run_once(p, seed, n_train=300, epochs=200, lr=1e-2):
    # TODO: torch.manual_seed(seed); data; modell; Adam; full-batch träning epochs steg; return {train_acc, test_acc, gap}
    ...


def ablation(ps=(0.0, 0.2, 0.5), seeds=(0, 1, 2), **kw):
    # TODO: {p: {"gap_mean", "gap_std", "test_mean", "test_std", "runs": [...]}}
    ...


def write_report(results, path="out/report.md"):
    # TODO: markdown med hypotes, metod (data, modell, budget, frön), tabell per p, slutsats
    ...

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

The gap (train−test) with p=0.5 is at least 15 % smaller than with p=0 (mean over 3 seeds, 300 epochs). At this small scale the effect is clear but smaller than in the paper — that is part of the reproduction's result and should be in the report. out/report.md reports mean ± std and seeds.

Common mistakes

  • Forgets model.eval() during evaluation → dropout active during testing.
  • The same seed for all runs → std = 0 and no uncertainty.
  • Measures test_acc on training data.