Images as matrices
Be able to manipulate an image as an array of pixel values and channels.
Prerequisites
Intuition
A colour image is a matrix of matrices: for every pixel, three numbers 0–255 (red, green, blue).
In code the image has a shape (height, width, channels). An image of 32 × 32 pixels in colour has the shape (32, 32, 3) and contains 3 072 numbers.
Doing things to the image means doing arithmetic on the numbers:
- Brighter: add 40 to every value (but clip at 255).
- Greyscale: merge the three channels into one value.
- Crop: take a rectangular cut-out of the matrix.
- Mirror: read the columns backwards.
Code
import numpy as np
image = np.zeros((4, 6, 3), dtype=np.uint8) # a black image, 4 rows × 6 columns
image[:, 3:, 0] = 255 # the right half red
print(image.shape, image.size) # (4, 6, 3) 72
grey = image.mean(axis=2) # merge the channels → (4, 6)
print(grey.shape, grey[0, 0], round(float(grey[0, 4]), 1)) # (4, 6) 0.0 85.0
brighter = np.clip(image.astype(int) + 40, 0, 255).astype(np.uint8)
cutout = image[1:3, 2:5] # rows 1–2, columns 2–4
mirrored = image[:, ::-1] # flip left/right
print(cutout.shape, mirrored.shape) # (2, 3, 3) (4, 6, 3)
A CNN does fundamentally the same thing: it slides a small filter across the matrix and does the arithmetic. The difference is that the filter's numbers are learnt instead of chosen by you.
Mastery means
- Describes an image as an array with height, width and channels
- Performs simple operations: greyscale, cropping, brightening
Sign in to do the exercises and build your mastery up.
Sources
- NumPy — dokumentation (BSD-3) — BSD-3-Clause
- CS231n — Image Classification — free to read