Skip to content
AI-grafen
EUniversityComputer vision· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

CNN architectures: LeNet to ResNet

Be able to explain how the architectures developed and why ResNet became the standard.

Prerequisites

Intuition

The history of CNN architectures is a series of answers to concrete problems.

YearArchitectureThe noveltyThe problem it solved
1998LeNet-5convolution + pooling, 7 layershandwritten digits
2012AlexNetReLU, dropout, GPU, augmentationImageNet — halved the error
2014VGG3×3 kernels only, 16–19 layerssimplicity and depth
2014Inceptionparallel kernel sizes, 1×1computational efficiency
2015ResNetresidual connections, 152 layersdepth itself
2017MobileNetdepthwise separable convolutionmobile hardware
2019EfficientNetbalanced scaling of depth/width/resolutionthe best quality per FLOP
2022ConvNeXta CNN with the transformer recipecompeting with ViT

Two insights recur:

  1. Small kernels stacked beat large ones. Two 3×3 layers have the same receptive field as one 5×5 but fewer parameters (18 against 25 per channel pair) and two non-linearities instead of one.
  2. A 1×1 convolution is free dimension reduction. It mixes the channels without touching the spatial dimensions, and is used to shrink the channel count before expensive operations.

Formal

Why ResNet was the turning point. Before 2015 depth was a problem: a 56-layer network performed worse than a 20-layer one, and not because of overfitting but because of optimisation. The residual connection y=x+F(x)y = x + F(x) made the identity mapping trivial to learn and gave the gradient an unscaled path through the network.

The effect was immediate: 152 layers, better than everything before it, and the principle spread to essentially every deep architecture built since — including the transformer.

The computational cost per layer:

FLOPs≈Hout⋅Wout⋅Cin⋅Cout⋅k2\text{FLOPs} \approx H_{out} \cdot W_{out} \cdot C_{in} \cdot C_{out} \cdot k^2

That explains three design choices:

ChoiceThe effect on the formula
Downsample earlyH⋅WH\cdot W falls fourfold per step
1×1 before 3×3 (a bottleneck)CinC_{in} in the expensive term goes down
Depthwise separable convolutionCin⋅Cout⋅k2C_{in}\cdot C_{out}\cdot k^2 becomes Cin⋅k2+Cin⋅CoutC_{in}\cdot k^2 + C_{in}\cdot C_{out}

The last is MobileNet's idea and gives 8–9× fewer computations for a 3×3 convolution with many channels.

The position today:

NeedChoice
The default choice for imagesa pretrained ResNet or ConvNeXt, fine-tuned
A small dataseta pretrained model — never train from scratch
Mobile or embeddedMobileNet, EfficientNet-Lite
Very large dataViT or a hybrid
Segmentation, detectiona backbone plus a task-specific head

The most important practical conclusion is that the choice of architecture is rarely what decides. Pretraining, data augmentation and data quality nearly always give more than swapping one modern architecture for another — the difference between ResNet-50 and ConvNeXt-T on a typical fine-tuning problem is often smaller than the difference between good and bad augmentation.

Code

import torch, torch.nn as nn

# 1. Two 3×3 beat one 5×5: the same receptive field, fewer parameters, more non-linearities
C = 64
one_5x5 = nn.Conv2d(C, C, 5, padding=2, bias=False)
two_3x3 = nn.Sequential(nn.Conv2d(C, C, 3, padding=1, bias=False), nn.ReLU(),
                        nn.Conv2d(C, C, 3, padding=1, bias=False))
for name, m in (("one 5×5", one_5x5), ("two 3×3", two_3x3)):
    print(f"{name:<9} {sum(p.numel() for p in m.parameters()):>8,} parameters")
# one 5×5    102,400 parameters
# two 3×3     73,728 parameters

# 2. The bottleneck block: 1×1 down, 3×3, 1×1 up
class Bottleneck(nn.Module):
    def __init__(self, channels, shrink=4):
        super().__init__()
        m = channels // shrink
        self.f = nn.Sequential(
            nn.Conv2d(channels, m, 1, bias=False), nn.BatchNorm2d(m), nn.ReLU(),
            nn.Conv2d(m, m, 3, padding=1, bias=False), nn.BatchNorm2d(m), nn.ReLU(),
            nn.Conv2d(m, channels, 1, bias=False), nn.BatchNorm2d(channels),
        )
        nn.init.zeros_(self.f[-1].weight)          # starts as the identity

    def forward(self, x):
        return torch.relu(x + self.f(x))

class Plain(nn.Module):
    def __init__(self, channels):
        super().__init__()
        self.f = nn.Sequential(
            nn.Conv2d(channels, channels, 3, padding=1, bias=False), nn.BatchNorm2d(channels),
            nn.ReLU(),
            nn.Conv2d(channels, channels, 3, padding=1, bias=False), nn.BatchNorm2d(channels),
        )

    def forward(self, x):
        return torch.relu(x + self.f(x))

for name, m in (("plain block", Plain(256)), ("bottleneck", Bottleneck(256))):
    print(f"{name:<14} {sum(p.numel() for p in m.parameters()):>9,} parameters")
# plain block      1,180,672 parameters
# bottleneck          70,400 parameters     ← 17× fewer, the same depth

# 3. Depthwise separable convolution (MobileNet)
def separable(cin, cout, k=3):
    return nn.Sequential(
        nn.Conv2d(cin, cin, k, padding=k // 2, groups=cin, bias=False),   # per channel
        nn.BatchNorm2d(cin), nn.ReLU(),
        nn.Conv2d(cin, cout, 1, bias=False),                              # mix the channels
        nn.BatchNorm2d(cout), nn.ReLU(),
    )

ordinary = nn.Conv2d(128, 256, 3, padding=1, bias=False)
sep = separable(128, 256)
print(f"ordinary 3×3: {sum(p.numel() for p in ordinary.parameters()):>8,}")
print(f"separable:    {sum(p.numel() for p in sep.parameters()):>8,}")
# ordinary 3×3:  294,912
# separable:      34,688     ← ~8× fewer

# 4. In practice: use a pretrained model
from torchvision.models import resnet50, ResNet50_Weights
m = resnet50(weights=ResNet50_Weights.IMAGENET1K_V2)
m.fc = nn.Linear(m.fc.in_features, 10)      # swap the head for your classes

Mastery means

  • Describes the development from LeNet to ResNet
  • Explains why each step was an improvement
  • Knows what applies today

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences