CNN architectures: LeNet to ResNet
Be able to explain how the architectures developed and why ResNet became the standard.
Prerequisites
- EConvolutional networks (CNNs)required
- EResidual connectionsrequired
Intuition
The history of CNN architectures is a series of answers to concrete problems.
| Year | Architecture | The novelty | The problem it solved |
|---|---|---|---|
| 1998 | LeNet-5 | convolution + pooling, 7 layers | handwritten digits |
| 2012 | AlexNet | ReLU, dropout, GPU, augmentation | ImageNet — halved the error |
| 2014 | VGG | 3×3 kernels only, 16–19 layers | simplicity and depth |
| 2014 | Inception | parallel kernel sizes, 1×1 | computational efficiency |
| 2015 | ResNet | residual connections, 152 layers | depth itself |
| 2017 | MobileNet | depthwise separable convolution | mobile hardware |
| 2019 | EfficientNet | balanced scaling of depth/width/resolution | the best quality per FLOP |
| 2022 | ConvNeXt | a CNN with the transformer recipe | competing with ViT |
Two insights recur:
- Small kernels stacked beat large ones. Two 3×3 layers have the same receptive field as one 5×5 but fewer parameters (18 against 25 per channel pair) and two non-linearities instead of one.
- A 1×1 convolution is free dimension reduction. It mixes the channels without touching the spatial dimensions, and is used to shrink the channel count before expensive operations.
Formal
Why ResNet was the turning point. Before 2015 depth was a problem: a 56-layer network performed worse than a 20-layer one, and not because of overfitting but because of optimisation. The residual connection made the identity mapping trivial to learn and gave the gradient an unscaled path through the network.
The effect was immediate: 152 layers, better than everything before it, and the principle spread to essentially every deep architecture built since — including the transformer.
The computational cost per layer:
That explains three design choices:
| Choice | The effect on the formula |
|---|---|
| Downsample early | falls fourfold per step |
| 1×1 before 3×3 (a bottleneck) | in the expensive term goes down |
| Depthwise separable convolution | becomes |
The last is MobileNet's idea and gives 8–9× fewer computations for a 3×3 convolution with many channels.
The position today:
| Need | Choice |
|---|---|
| The default choice for images | a pretrained ResNet or ConvNeXt, fine-tuned |
| A small dataset | a pretrained model — never train from scratch |
| Mobile or embedded | MobileNet, EfficientNet-Lite |
| Very large data | ViT or a hybrid |
| Segmentation, detection | a backbone plus a task-specific head |
The most important practical conclusion is that the choice of architecture is rarely what decides. Pretraining, data augmentation and data quality nearly always give more than swapping one modern architecture for another — the difference between ResNet-50 and ConvNeXt-T on a typical fine-tuning problem is often smaller than the difference between good and bad augmentation.
Code
import torch, torch.nn as nn
# 1. Two 3×3 beat one 5×5: the same receptive field, fewer parameters, more non-linearities
C = 64
one_5x5 = nn.Conv2d(C, C, 5, padding=2, bias=False)
two_3x3 = nn.Sequential(nn.Conv2d(C, C, 3, padding=1, bias=False), nn.ReLU(),
nn.Conv2d(C, C, 3, padding=1, bias=False))
for name, m in (("one 5×5", one_5x5), ("two 3×3", two_3x3)):
print(f"{name:<9} {sum(p.numel() for p in m.parameters()):>8,} parameters")
# one 5×5 102,400 parameters
# two 3×3 73,728 parameters
# 2. The bottleneck block: 1×1 down, 3×3, 1×1 up
class Bottleneck(nn.Module):
def __init__(self, channels, shrink=4):
super().__init__()
m = channels // shrink
self.f = nn.Sequential(
nn.Conv2d(channels, m, 1, bias=False), nn.BatchNorm2d(m), nn.ReLU(),
nn.Conv2d(m, m, 3, padding=1, bias=False), nn.BatchNorm2d(m), nn.ReLU(),
nn.Conv2d(m, channels, 1, bias=False), nn.BatchNorm2d(channels),
)
nn.init.zeros_(self.f[-1].weight) # starts as the identity
def forward(self, x):
return torch.relu(x + self.f(x))
class Plain(nn.Module):
def __init__(self, channels):
super().__init__()
self.f = nn.Sequential(
nn.Conv2d(channels, channels, 3, padding=1, bias=False), nn.BatchNorm2d(channels),
nn.ReLU(),
nn.Conv2d(channels, channels, 3, padding=1, bias=False), nn.BatchNorm2d(channels),
)
def forward(self, x):
return torch.relu(x + self.f(x))
for name, m in (("plain block", Plain(256)), ("bottleneck", Bottleneck(256))):
print(f"{name:<14} {sum(p.numel() for p in m.parameters()):>9,} parameters")
# plain block 1,180,672 parameters
# bottleneck 70,400 parameters ← 17× fewer, the same depth
# 3. Depthwise separable convolution (MobileNet)
def separable(cin, cout, k=3):
return nn.Sequential(
nn.Conv2d(cin, cin, k, padding=k // 2, groups=cin, bias=False), # per channel
nn.BatchNorm2d(cin), nn.ReLU(),
nn.Conv2d(cin, cout, 1, bias=False), # mix the channels
nn.BatchNorm2d(cout), nn.ReLU(),
)
ordinary = nn.Conv2d(128, 256, 3, padding=1, bias=False)
sep = separable(128, 256)
print(f"ordinary 3×3: {sum(p.numel() for p in ordinary.parameters()):>8,}")
print(f"separable: {sum(p.numel() for p in sep.parameters()):>8,}")
# ordinary 3×3: 294,912
# separable: 34,688 ← ~8× fewer
# 4. In practice: use a pretrained model
from torchvision.models import resnet50, ResNet50_Weights
m = resnet50(weights=ResNet50_Weights.IMAGENET1K_V2)
m.fc = nn.Linear(m.fc.in_features, 10) # swap the head for your classes
Mastery means
- Describes the development from LeNet to ResNet
- Explains why each step was an improvement
- Knows what applies today
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Deep Residual Learning for Image Recognition — arXiv (open access; licence per article)
- arXiv — A ConvNet for the 2020s (ConvNeXt) — arXiv (open access; licence per article)
- PyTorch — tutorials (BSD-3) — BSD-3-Clause