Skip to content
AI-grafen
EUniversityComputer vision· about 60 min· fundamentals that rarely change· verified 2026-09-20· EN

Convolutional networks (CNNs)

Be able to explain convolution, pooling and receptive fields, and build a CNN for image classification.

Prerequisites

Intuition

An image of 224 × 224 × 3 has 150 000 numbers. A fully connected layer of 1 000 neurons would have 150 million weights — and would learn the same edge separately in a thousand different places.

Convolution: a small filter (3 × 3 × 3 = 27 weights) slides across the image and computes the same thing everywhere. 64 filters → 64 «maps» of where something is (edges, colour transitions). The same weights everywhere = translation invariance + few parameters.

Pooling (max or mean over 2 × 2) halves the resolution; after a few layers each neuron covers a large receptive field and sees whole objects. Stack them: convolution → ReLU → pooling, repeat, and a small fully connected head at the end.

Modern alternatives (ViT) do something similar with attention over image patches, but CNNs are still the most efficient on small data.

Code

import torch, torch.nn as nn

class SmallCNN(nn.Module):
    def __init__(self, classes=10):
        super().__init__()
        self.f = nn.Sequential(
            nn.Conv2d(3, 32, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),   # 32×32 → 16×16
            nn.Conv2d(32, 64, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),  # → 8×8
            nn.Conv2d(64, 128, 3, padding=1), nn.ReLU(), nn.AdaptiveAvgPool2d(1),
        )
        self.head = nn.Linear(128, classes)
    def forward(self, x):
        return self.head(self.f(x).flatten(1))

m = SmallCNN()
print(sum(p.numel() for p in m.parameters()))    # ≈ 94 000 — compare 150 M for a dense layer
print(m(torch.zeros(4, 3, 32, 32)).shape)         # (4, 10)

Parameters in a Conv2d(in, out, k): in · out · k · k + out. Output size with padding p and stride s: ⌊(H + 2p − k)/s⌋ + 1.

Formal

2D convolution (really cross-correlation in DL libraries): (x∗w)[i,j]=∑c∑u,vx[c,i+u,j+v] w[c,u,v](x * w)[i,j] = \sum_{c}\sum_{u,v} x[c, i+u, j+v]\, w[c,u,v]. A layer with CinC_{in} input and CoutC_{out} output channels and kernel kk has Cout(Cink2+1)C_{out}(C_{in}k^2 + 1) parameters, independent of the image size. The receptive field grows additively: after LL layers with k=3k=3, stride 1: 1+2L1 + 2L; every pooling or stride-2 doubles the growth. Equivariance: f(Tδx)=Tδf(x)f(T_\delta x) = T_\delta f(x) for translations δ\delta (up to edge effects); pooling gives approximate invariance. Batch normalisation and residual connections are what make deep CNNs (ResNet) trainable.

Mastery means

  • Explains convolution, pooling and receptive fields
  • Calculates the output size and the parameters of a convolutional layer
  • Builds and trains a small CNN for image classification

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences