Residual connections
Be able to explain why skip connections make deep networks trainable.
Prerequisites
- EVanishing and exploding gradientsrequired
Intuition
Before 2015 it was not possible to train really deep networks. A 56-layer network was worse than a 20-layer one — not because of overfitting, but because it could not be optimised. That was a puzzling result: the deeper network could in principle imitate the shallower one by letting the extra layers do nothing.
The problem was that «doing nothing» — the identity mapping — is surprisingly hard to learn with a stack of weight matrices.
The residual connection's solution is to build it in:
The layer learns the difference rather than the whole mapping. If the layer should do nothing, it is enough that , which is trivial: set the weights close to zero.
The result was immediate. ResNet in 2015 trained 152 layers where earlier networks choked at 20.
Formal
The gradient flow is the mathematical explanation. With :
Through layers the gradient becomes a product of such terms. The one in every factor guarantees a path where the gradient is not scaled down. Without the residual it is a product of Jacobians, and if each is on average smaller than 1 the gradient dies exponentially with depth.
A residual stack can also be seen as an ensemble of shallow paths: there are different ways through the network (take or skip each block), and most of the effective paths are short. That explains why the network works even if you remove a whole block — something that knocks out an ordinary deep network completely.
Pre-norm versus post-norm — a detail that decides whether large transformers can be trained at all:
| Variant | Formula | Property |
|---|---|---|
| Post-norm (the original) | a better final result, but it needs warm-up and is sensitive | |
| Pre-norm | stable, trains without warm-up, the standard today |
The difference: in pre-norm there is a clean, unnormalised path from the input to the output. In post-norm the residual passes through LayerNorm in every layer, which rescales it and breaks the straight gradient path. At 50+ layers the difference is decisive, and that is why essentially every modern language model uses pre-norm.
Two practical details:
- Dimension matching. If the block changes the number of channels, the shortcut needs a projection — a 1×1 convolution or a linear layer.
- Zero-initialising the last layer in the block makes the block start as exactly the identity. That noticeably stabilises the start of training and is used in many modern recipes.
Code
import torch, torch.nn as nn
class ResidualBlock(nn.Module):
def __init__(self, channels):
super().__init__()
self.f = nn.Sequential(
nn.Conv2d(channels, channels, 3, padding=1, bias=False),
nn.BatchNorm2d(channels), nn.ReLU(),
nn.Conv2d(channels, channels, 3, padding=1, bias=False),
nn.BatchNorm2d(channels),
)
nn.init.zeros_(self.f[-1].weight) # the block starts as exactly the identity
def forward(self, x):
return torch.relu(x + self.f(x))
# The pre-norm block in a transformer
class PreNormBlock(nn.Module):
def __init__(self, d, heads):
super().__init__()
self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
self.attn = nn.MultiheadAttention(d, heads, batch_first=True)
self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))
def forward(self, x):
x = x + self.attn(self.n1(x), self.n1(x), self.n1(x), need_weights=False)[0]
return x + self.mlp(self.n2(x)) # a clean path through x, unchanged by the LN
# Measure the gradient flow: with and without a residual
def gradient_in_first_layer(with_residual, depth=40, d=64):
layers = nn.ModuleList([nn.Sequential(nn.Linear(d, d), nn.Tanh()) for _ in range(depth)])
x = torch.randn(8, d, requires_grad=True)
h = x
for lg in layers:
h = h + lg(h) if with_residual else lg(h)
h.sum().backward()
return float(x.grad.norm())
torch.manual_seed(0)
print(f"without a residual: {gradient_in_first_layer(False):.3e}")
print(f"with a residual: {gradient_in_first_layer(True):.3e}")
# without a residual: 1.4e-06 ← the gradient has practically died on the way
# with a residual: 2.6e+01 ← it arrives
The difference in the last output is seven orders of magnitude through forty layers. That is exactly what made deep networks possible, and nn.init.zeros_ on the block's last layer is one line that stabilises the first hundred steps of training for free.
Mastery means
- Explains how a residual connection affects the gradient flow
- Knows the difference between pre-norm and post-norm
- Knows why the identity mapping is hard to learn
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Deep Residual Learning for Image Recognition — arXiv (open access; licence per article)
- arXiv — On Layer Normalization in the Transformer Architecture — arXiv (open access; licence per article)
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0