Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Dropout in detail

Be able to explain dropout as an ensemble and why it is turned off at inference.

Prerequisites

Intuition

During training every activation is zeroed with probability p, independently. Every minibatch therefore trains a random subnetwork.

Two ways of understanding why that helps:

  1. An implicit ensemble. With n neurons there are 2ⁿ possible subnetworks. The training trains them together with shared weights, and at inference (without dropout) the average of all of them is approximated. Ensembles reduce variance — that is precisely the effect.
  2. It counteracts co-adaptation. A neuron cannot rely on a particular other neuron still being there, so it has to contribute something useful of its own.

At inference dropout is turned off — otherwise the answer becomes random. That is why model.eval() is not optional.

Formal

Inverted dropout (what every library uses) scales during training instead of at inference:

atrain=m⊙a1−p,mi∼Bernoulli(1−p)a_{\text{train}} = \frac{m \odot a}{1-p}, \qquad m_i \sim \text{Bernoulli}(1-p)

Then E[atrain]=a\mathbb E[a_{\text{train}}] = a, and the inference does not have to do anything at all — which is the point: the inference code becomes identical with and without dropout.

(The original formulation instead scaled by (1−p)(1-p) at inference. Mathematically equivalent, practically worse.)

Choosing p:

The placementA typical p
Fully connected layers0.3–0.5
Convolutional layers0.0–0.2 (few parameters, strong weight sharing)
Transformer blocks (residual, attention)0.0–0.1
The input layer0.1–0.2 if at all

Large transformers often use no dropout at all in pre-training: with enough data overfitting is not the problem, and dropout then only costs effective capacity. In fine-tuning on little data it comes back.

MC dropout: leave dropout on at inference and sample n times. The spread in the predictions becomes a rough uncertainty estimate. Cheap but blunt — better than nothing when calibrated uncertainty is needed.

Code

import numpy as np

def dropout_forward(a, p, training, rng):
    if not training or p == 0:
        return a                                  # inference: nothing happens
    mask = (rng.random(a.shape) > p).astype(a.dtype)
    return a * mask / (1 - p)                     # inverted: scale during TRAINING

rng = np.random.default_rng(0)
a = np.ones(10_000)
out = dropout_forward(a, p=0.3, training=True, rng=rng)
print(round(out.mean(), 3), round((out == 0).mean(), 3))  # 1.002 0.299
#     the expectation preserved ↑        ↑ ~30 % zeroed

# MC dropout for uncertainty
def mc_dropout(model, x, n=30):
    model.train()                                 # dropout ON, deliberately
    with torch.no_grad():
        p = torch.stack([model(x).softmax(-1) for _ in range(n)])
    model.eval()
    return p.mean(0), p.std(0)                    # the mean and the uncertainty per class

Mastery means

  • Explains dropout as an implicit ensemble
  • Describes inverted dropout and why it is turned off at inference
  • Chooses p according to the type of layer

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences