Skip to content
AI-grafen
DAI developerComputer vision· about 45 min· fundamentals that rarely change· verified 2026-09-20· EN

Convolution by hand

Be able to carry out a 2D convolution by hand and explain edge detection with filters.

Prerequisites

Intuition

A convolution is sliding a small grid of numbers — a kernel — across the image and, at each position, computing a weighted sum.

Image (5×5)             Kernel (3×3)
0  0  0  0  0           -1  0  1
0 10 10 10  0           -1  0  1      ← responds to vertical edges
0 10 10 10  0           -1  0  1
0 10 10 10  0
0  0  0  0  0

Lay the kernel over the top left 3×3 area, multiply element by element, sum. Move one step to the right. Repeat.

The calculation for position (0,0) — the area is

 0  0  0
 0 10 10
 0 10 10

The sum = (−1·0 + 0·0 + 1·0) + (−1·0 + 0·10 + 1·10) + (−1·0 + 0·10 + 1·10) = 20

A positive value means «dark on the left, light on the right» — an edge. That is the whole idea: a small matrix of numbers becomes an edge detector.

Formal

The formula (really cross-correlation, which is what every framework calls «convolution»):

yij=∑u=0k−1∑v=0k−1xi+u, j+v⋅wuvy_{ij} = \sum_{u=0}^{k-1}\sum_{v=0}^{k-1} x_{i+u,\,j+v} \cdot w_{uv}

The size of the output:

Hout=⌊Hin+2p−ks⌋+1H_{out} = \left\lfloor\frac{H_{in} + 2p - k}{s}\right\rfloor + 1

where pp is the padding and ss is the stride.

HinH_{in}kkppssHoutH_{out}
53013
53115 (unchanged)
2831214 (halved)
3252132

padding = (k−1)/2 with stride = 1 preserves the size — which is why odd kernel sizes are the standard.

Classic kernels:

KernelMatrixDoes
Identity0 0 0 / 0 1 0 / 0 0 0nothing
Sobel vertical-1 0 1 / -2 0 2 / -1 0 1vertical edges
Sobel horizontal-1 -2 -1 / 0 0 0 / 1 2 1horizontal edges
Blur1/9 · all onessmooths
Sharpen0 -1 0 / -1 5 -1 / 0 -1 0enhances the detail
Laplacian0 1 0 / 1 -4 1 / 0 1 0edges in every direction

Two properties that make convolution so effective in networks:

  1. Parameter sharing. The same nine numbers are used across the whole image. A fully connected layer between two 28×28 images would need 614 656 weights; a 3×3 kernel needs 9.
  2. Locality. Every output value depends only on a small area — which matches how images actually work.

The big difference in a CNN: the kernels are not hand-written. They are learnt. And the curious thing is that the first layers of a trained network nearly always learn something resembling Sobel filters — edge detectors, entirely by themselves.

Code

import numpy as np

def convolve(image, kernel, padding=0, stride=1):
    if padding:
        image = np.pad(image, padding)
    k = kernel.shape[0]
    H = (image.shape[0] - k) // stride + 1
    W = (image.shape[1] - k) // stride + 1
    out = np.zeros((H, W))
    for i in range(H):
        for j in range(W):
            area = image[i*stride:i*stride+k, j*stride:j*stride+k]
            out[i, j] = float((area * kernel).sum())
    return out

image = np.zeros((5, 5)); image[1:4, 1:4] = 10
vertical = np.array([[-1, 0, 1], [-1, 0, 1], [-1, 0, 1]])

print(convolve(image, vertical))
# [[ 20.   0. -20.]
#  [ 30.   0. -30.]
#  [ 20.   0. -20.]]
#  ↑ positive at the left edge, negative at the right edge, zero in the middle

# The size formula
for H, k, p, s in [(5, 3, 0, 1), (5, 3, 1, 1), (28, 3, 1, 2), (32, 5, 2, 1)]:
    print(f"H={H} k={k} p={p} s={s} → {(H + 2*p - k)//s + 1}")
# H=5 k=3 p=0 s=1 → 3
# H=5 k=3 p=1 s=1 → 5     ← padding preserves the size
# H=28 k=3 p=1 s=2 → 14   ← the stride halves it
# H=32 k=5 p=2 s=1 → 32

# Parameter sharing — why CNNs work
print(28*28 * 28*28, "weights in a fully connected layer")   # 614656
print(3*3, "weights in a 3×3 kernel")                        # 9

Mastery means

  • Works out a convolution by hand
  • Explains what different kernels do
  • Works out the size of the output

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences