Dropout in detail
Be able to explain dropout as an ensemble and why it is turned off at inference.
Prerequisites
Intuition
During training every activation is zeroed with probability p, independently. Every minibatch therefore trains a random subnetwork.
Two ways of understanding why that helps:
- An implicit ensemble. With n neurons there are 2ⁿ possible subnetworks. The training trains them together with shared weights, and at inference (without dropout) the average of all of them is approximated. Ensembles reduce variance — that is precisely the effect.
- It counteracts co-adaptation. A neuron cannot rely on a particular other neuron still being there, so it has to contribute something useful of its own.
At inference dropout is turned off — otherwise the answer becomes random. That is why model.eval() is not optional.
Formal
Inverted dropout (what every library uses) scales during training instead of at inference:
Then , and the inference does not have to do anything at all — which is the point: the inference code becomes identical with and without dropout.
(The original formulation instead scaled by at inference. Mathematically equivalent, practically worse.)
Choosing p:
| The placement | A typical p |
|---|---|
| Fully connected layers | 0.3–0.5 |
| Convolutional layers | 0.0–0.2 (few parameters, strong weight sharing) |
| Transformer blocks (residual, attention) | 0.0–0.1 |
| The input layer | 0.1–0.2 if at all |
Large transformers often use no dropout at all in pre-training: with enough data overfitting is not the problem, and dropout then only costs effective capacity. In fine-tuning on little data it comes back.
MC dropout: leave dropout on at inference and sample n times. The spread in the predictions becomes a rough uncertainty estimate. Cheap but blunt — better than nothing when calibrated uncertainty is needed.
Code
import numpy as np
def dropout_forward(a, p, training, rng):
if not training or p == 0:
return a # inference: nothing happens
mask = (rng.random(a.shape) > p).astype(a.dtype)
return a * mask / (1 - p) # inverted: scale during TRAINING
rng = np.random.default_rng(0)
a = np.ones(10_000)
out = dropout_forward(a, p=0.3, training=True, rng=rng)
print(round(out.mean(), 3), round((out == 0).mean(), 3)) # 1.002 0.299
# the expectation preserved ↑ ↑ ~30 % zeroed
# MC dropout for uncertainty
def mc_dropout(model, x, n=30):
model.train() # dropout ON, deliberately
with torch.no_grad():
p = torch.stack([model(x).softmax(-1) for _ in range(n)])
model.eval()
return p.mean(0), p.std(0) # the mean and the uncertainty per class
Mastery means
- Explains dropout as an implicit ensemble
- Describes inverted dropout and why it is turned off at inference
- Chooses p according to the type of layer
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR) — arXiv (open access; licence per article)
- arXiv — Dropout as a Bayesian Approximation (MC-dropout) — arXiv (open access; licence per article)