Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Residual connections

Be able to explain why skip connections make deep networks trainable.

Prerequisites

Intuition

Before 2015 it was not possible to train really deep networks. A 56-layer network was worse than a 20-layer one — not because of overfitting, but because it could not be optimised. That was a puzzling result: the deeper network could in principle imitate the shallower one by letting the extra layers do nothing.

The problem was that «doing nothing» — the identity mapping — is surprisingly hard to learn with a stack of weight matrices.

The residual connection's solution is to build it in:

y=x+F(x)y = x + F(x)

The layer learns the difference F(x)F(x) rather than the whole mapping. If the layer should do nothing, it is enough that F(x)→0F(x) \to 0, which is trivial: set the weights close to zero.

The result was immediate. ResNet in 2015 trained 152 layers where earlier networks choked at 20.

Formal

The gradient flow is the mathematical explanation. With y=x+F(x)y = x + F(x):

∂y∂x=I+∂F∂x\frac{\partial y}{\partial x} = I + \frac{\partial F}{\partial x}

Through LL layers the gradient becomes a product of such terms. The one in every factor guarantees a path where the gradient is not scaled down. Without the residual it is a product of LL Jacobians, and if each is on average smaller than 1 the gradient dies exponentially with depth.

A residual stack can also be seen as an ensemble of shallow paths: there are 2L2^L different ways through the network (take or skip each block), and most of the effective paths are short. That explains why the network works even if you remove a whole block — something that knocks out an ordinary deep network completely.

Pre-norm versus post-norm — a detail that decides whether large transformers can be trained at all:

VariantFormulaProperty
Post-norm (the original)y=LN(x+F(x))y = \mathrm{LN}(x + F(x))a better final result, but it needs warm-up and is sensitive
Pre-normy=x+F(LN(x))y = x + F(\mathrm{LN}(x))stable, trains without warm-up, the standard today

The difference: in pre-norm there is a clean, unnormalised path from the input to the output. In post-norm the residual passes through LayerNorm in every layer, which rescales it and breaks the straight gradient path. At 50+ layers the difference is decisive, and that is why essentially every modern language model uses pre-norm.

Two practical details:

  1. Dimension matching. If the block changes the number of channels, the shortcut needs a projection — a 1×1 convolution or a linear layer.
  2. Zero-initialising the last layer in the block makes the block start as exactly the identity. That noticeably stabilises the start of training and is used in many modern recipes.

Code

import torch, torch.nn as nn

class ResidualBlock(nn.Module):
    def __init__(self, channels):
        super().__init__()
        self.f = nn.Sequential(
            nn.Conv2d(channels, channels, 3, padding=1, bias=False),
            nn.BatchNorm2d(channels), nn.ReLU(),
            nn.Conv2d(channels, channels, 3, padding=1, bias=False),
            nn.BatchNorm2d(channels),
        )
        nn.init.zeros_(self.f[-1].weight)     # the block starts as exactly the identity

    def forward(self, x):
        return torch.relu(x + self.f(x))

# The pre-norm block in a transformer
class PreNormBlock(nn.Module):
    def __init__(self, d, heads):
        super().__init__()
        self.n1, self.n2 = nn.LayerNorm(d), nn.LayerNorm(d)
        self.attn = nn.MultiheadAttention(d, heads, batch_first=True)
        self.mlp = nn.Sequential(nn.Linear(d, 4 * d), nn.GELU(), nn.Linear(4 * d, d))

    def forward(self, x):
        x = x + self.attn(self.n1(x), self.n1(x), self.n1(x), need_weights=False)[0]
        return x + self.mlp(self.n2(x))       # a clean path through x, unchanged by the LN

# Measure the gradient flow: with and without a residual
def gradient_in_first_layer(with_residual, depth=40, d=64):
    layers = nn.ModuleList([nn.Sequential(nn.Linear(d, d), nn.Tanh()) for _ in range(depth)])
    x = torch.randn(8, d, requires_grad=True)
    h = x
    for lg in layers:
        h = h + lg(h) if with_residual else lg(h)
    h.sum().backward()
    return float(x.grad.norm())

torch.manual_seed(0)
print(f"without a residual: {gradient_in_first_layer(False):.3e}")
print(f"with a residual:    {gradient_in_first_layer(True):.3e}")
# without a residual: 1.4e-06     ← the gradient has practically died on the way
# with a residual:    2.6e+01     ← it arrives

The difference in the last output is seven orders of magnitude through forty layers. That is exactly what made deep networks possible, and nn.init.zeros_ on the block's last layer is one line that stabilises the first hundred steps of training for free.

Mastery means

  • Explains how a residual connection affects the gradient flow
  • Knows the difference between pre-norm and post-norm
  • Knows why the identity mapping is hard to learn

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences