Convolution by hand
Be able to carry out a 2D convolution by hand and explain edge detection with filters.
Prerequisites
- CImages as matricesrequired
Intuition
A convolution is sliding a small grid of numbers — a kernel — across the image and, at each position, computing a weighted sum.
Image (5×5) Kernel (3×3)
0 0 0 0 0 -1 0 1
0 10 10 10 0 -1 0 1 ← responds to vertical edges
0 10 10 10 0 -1 0 1
0 10 10 10 0
0 0 0 0 0
Lay the kernel over the top left 3×3 area, multiply element by element, sum. Move one step to the right. Repeat.
The calculation for position (0,0) — the area is
0 0 0
0 10 10
0 10 10
The sum = (−1·0 + 0·0 + 1·0) + (−1·0 + 0·10 + 1·10) + (−1·0 + 0·10 + 1·10) = 20
A positive value means «dark on the left, light on the right» — an edge. That is the whole idea: a small matrix of numbers becomes an edge detector.
Formal
The formula (really cross-correlation, which is what every framework calls «convolution»):
The size of the output:
where is the padding and is the stride.
| 5 | 3 | 0 | 1 | 3 |
| 5 | 3 | 1 | 1 | 5 (unchanged) |
| 28 | 3 | 1 | 2 | 14 (halved) |
| 32 | 5 | 2 | 1 | 32 |
padding = (k−1)/2 with stride = 1 preserves the size — which is why odd kernel sizes are the standard.
Classic kernels:
| Kernel | Matrix | Does |
|---|---|---|
| Identity | 0 0 0 / 0 1 0 / 0 0 0 | nothing |
| Sobel vertical | -1 0 1 / -2 0 2 / -1 0 1 | vertical edges |
| Sobel horizontal | -1 -2 -1 / 0 0 0 / 1 2 1 | horizontal edges |
| Blur | 1/9 · all ones | smooths |
| Sharpen | 0 -1 0 / -1 5 -1 / 0 -1 0 | enhances the detail |
| Laplacian | 0 1 0 / 1 -4 1 / 0 1 0 | edges in every direction |
Two properties that make convolution so effective in networks:
- Parameter sharing. The same nine numbers are used across the whole image. A fully connected layer between two 28×28 images would need 614 656 weights; a 3×3 kernel needs 9.
- Locality. Every output value depends only on a small area — which matches how images actually work.
The big difference in a CNN: the kernels are not hand-written. They are learnt. And the curious thing is that the first layers of a trained network nearly always learn something resembling Sobel filters — edge detectors, entirely by themselves.
Code
import numpy as np
def convolve(image, kernel, padding=0, stride=1):
if padding:
image = np.pad(image, padding)
k = kernel.shape[0]
H = (image.shape[0] - k) // stride + 1
W = (image.shape[1] - k) // stride + 1
out = np.zeros((H, W))
for i in range(H):
for j in range(W):
area = image[i*stride:i*stride+k, j*stride:j*stride+k]
out[i, j] = float((area * kernel).sum())
return out
image = np.zeros((5, 5)); image[1:4, 1:4] = 10
vertical = np.array([[-1, 0, 1], [-1, 0, 1], [-1, 0, 1]])
print(convolve(image, vertical))
# [[ 20. 0. -20.]
# [ 30. 0. -30.]
# [ 20. 0. -20.]]
# ↑ positive at the left edge, negative at the right edge, zero in the middle
# The size formula
for H, k, p, s in [(5, 3, 0, 1), (5, 3, 1, 1), (28, 3, 1, 2), (32, 5, 2, 1)]:
print(f"H={H} k={k} p={p} s={s} → {(H + 2*p - k)//s + 1}")
# H=5 k=3 p=0 s=1 → 3
# H=5 k=3 p=1 s=1 → 5 ← padding preserves the size
# H=28 k=3 p=1 s=2 → 14 ← the stride halves it
# H=32 k=5 p=2 s=1 → 32
# Parameter sharing — why CNNs work
print(28*28 * 28*28, "weights in a fully connected layer") # 614656
print(3*3, "weights in a 3×3 kernel") # 9
Mastery means
- Works out a convolution by hand
- Explains what different kernels do
- Works out the size of the output
Sign in to do the exercises and build your mastery up.
Sources
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0
- CS Unplugged (CC BY-SA 4.0) — CC BY-SA 4.0
- Wikipedia — Kernel (image processing) (CC BY-SA 4.0) — CC BY-SA 4.0