Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Weight initialisation

Be able to explain why the initialisation matters and use Xavier/He.

Prerequisites

Intuition

All the weights zero? Then every neuron in a layer gives exactly the same output and the same gradient — they stay identical for ever. The network behaves as if it had a single neuron per layer. The symmetry has to be broken.

Weights too large? The activations grow layer by layer → the gradients explode → NaN.

Weights too small? The activations shrink towards zero → the gradients vanish → nothing learns.

The goal: keep the variance of the activations roughly constant through the whole network. That is precisely what Xavier and He initialisation are constructed for.

Formal

If you want Var(y)≈Var(x)\text{Var}(y) \approx \text{Var}(x) through a layer y=Wxy = Wx with ninn_{in} inputs, ninVar(w)=1n_{in}\text{Var}(w) = 1 is required.

Xavier/Glorot (for sigmoid/tanh, symmetric about zero): Var(w)=2nin+nout\text{Var}(w) = \dfrac{2}{n_{in}+n_{out}} — a compromise between the forward and the backward pass.

He/Kaiming (for ReLU): ReLU zeroes half the values and thereby halves the variance, so you compensate: Var(w)=2nin\text{Var}(w) = \dfrac{2}{n_{in}}.

The bias is normally initialised to 0.

Modern architectures are less sensitive thanks to normalisation layers (batch/layer norm) and residual connections, but the initialisation still decides the first hundred steps — and in very deep networks without normalisation it is the difference between convergence and NaN.

Code

import torch, torch.nn as nn

def variance_profile(init_fn, layers=8, width=512):
    x = torch.randn(1000, width)
    for _ in range(layers):
        lin = nn.Linear(width, width, bias=False)
        init_fn(lin.weight)
        x = torch.relu(lin(x))
    return x.var().item()

print("too small:", round(variance_profile(lambda w: nn.init.normal_(w, std=0.01)), 8))
print("He:       ", round(variance_profile(lambda w: nn.init.kaiming_normal_(w, nonlinearity="relu")), 4))
print("too large:", round(variance_profile(lambda w: nn.init.normal_(w, std=0.5)), 1))
# too small:  0.00000000   ← the activations have died out
# He:         ~0.3         ← the same order of magnitude through the whole net
# too large:  1e+08        ← it explodes

In PyTorch nn.Linear and nn.Conv2d are initialised with a variant of Kaiming from the start — but when you write your own layers or load weights manually it is your responsibility.

Mastery means

  • Explains why the initialisation affects the training
  • Chooses Xavier or He according to the activation function

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences