Skip to content
AI-grafen
EUniversityDeep learning· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Batch and layer normalisation

Be able to explain what normalisation does to the activations and when to choose batch or layer norm.

Prerequisites

Intuition

A normalisation layer takes the activations, subtracts the mean and divides by the standard deviation — and then learns two parameters (γ, β) that can scale and shift back if needed.

Batch norm normalises over the batch for each feature. It therefore needs a batch, and behaves differently in training (the batch's statistics) and at inference (running averages). The standard in CNNs.

Layer norm normalises over the features within each example. Independent of the batch size, identical in training and at inference. The standard in transformers — among other reasons because sequences have different lengths and batch statistics become unreliable.

The effect: more stable training, a higher learning rate becomes usable, less sensitivity to the initialisation.

Code

import torch, torch.nn as nn

x = torch.randn(32, 64)                    # (batch, features)

bn = nn.BatchNorm1d(64)
ln = nn.LayerNorm(64)

print(bn(x).mean(dim=0)[:3].round(decimals=4))   # ≈ 0 per feature (over the batch)
print(ln(x).mean(dim=1)[:3].round(decimals=4))   # ≈ 0 per example (over the features)

# Batch norm behaves differently in eval:
bn.eval()
print(bn(x).mean().item())      # uses the running statistics, not the batch's

# A transformer block: layer norm before the attention (pre-norm) is the standard today
class Block(nn.Module):
    def __init__(self, d):
        super().__init__()
        self.ln1, self.ln2 = nn.LayerNorm(d), nn.LayerNorm(d)
        self.attn, self.mlp = nn.MultiheadAttention(d, 8, batch_first=True), nn.Sequential(nn.Linear(d, 4*d), nn.GELU(), nn.Linear(4*d, d))
    def forward(self, x):
        h = self.ln1(x); x = x + self.attn(h, h, h, need_weights=False)[0]
        return x + self.mlp(self.ln2(x))

The trap: batch norm with a batch size of 1 or 2 gives meaningless statistics and can destroy the training completely. Use layer norm or GroupNorm when the batch has to be small.

Mastery means

  • Explains what normalisation does to the activations
  • Chooses batch norm or layer norm according to the situation

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences