Weight initialisation
Be able to explain why the initialisation matters and use Xavier/He.
Prerequisites
Intuition
All the weights zero? Then every neuron in a layer gives exactly the same output and the same gradient — they stay identical for ever. The network behaves as if it had a single neuron per layer. The symmetry has to be broken.
Weights too large? The activations grow layer by layer → the gradients explode → NaN.
Weights too small? The activations shrink towards zero → the gradients vanish → nothing learns.
The goal: keep the variance of the activations roughly constant through the whole network. That is precisely what Xavier and He initialisation are constructed for.
Formal
If you want through a layer with inputs, is required.
Xavier/Glorot (for sigmoid/tanh, symmetric about zero): — a compromise between the forward and the backward pass.
He/Kaiming (for ReLU): ReLU zeroes half the values and thereby halves the variance, so you compensate: .
The bias is normally initialised to 0.
Modern architectures are less sensitive thanks to normalisation layers (batch/layer norm) and residual connections, but the initialisation still decides the first hundred steps — and in very deep networks without normalisation it is the difference between convergence and NaN.
Code
import torch, torch.nn as nn
def variance_profile(init_fn, layers=8, width=512):
x = torch.randn(1000, width)
for _ in range(layers):
lin = nn.Linear(width, width, bias=False)
init_fn(lin.weight)
x = torch.relu(lin(x))
return x.var().item()
print("too small:", round(variance_profile(lambda w: nn.init.normal_(w, std=0.01)), 8))
print("He: ", round(variance_profile(lambda w: nn.init.kaiming_normal_(w, nonlinearity="relu")), 4))
print("too large:", round(variance_profile(lambda w: nn.init.normal_(w, std=0.5)), 1))
# too small: 0.00000000 ← the activations have died out
# He: ~0.3 ← the same order of magnitude through the whole net
# too large: 1e+08 ← it explodes
In PyTorch nn.Linear and nn.Conv2d are initialised with a variant of Kaiming from the start — but when you write your own layers or load weights manually it is your responsibility.
Mastery means
- Explains why the initialisation affects the training
- Chooses Xavier or He according to the activation function
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Delving Deep into Rectifiers (He-init) — arXiv (open access; licence per article)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0