Image preprocessing and augmentation
Be able to load, scale, normalise and augment images.
Prerequisites
Intuition
Before an image reaches the model it goes through four steps, in this order:
- Load and decode (JPEG → an array). Check the channel order: OpenCV gives BGR, PIL gives RGB. That is a classic and silent bug.
- Scale it to the size the model expects. Preserve the proportions if the shape means something — otherwise objects come out stretched.
- Normalise with the same mean and standard deviation the pretraining used.
- Augment — but only on the training data.
Step 3 is the one most often forgotten in transfer learning, and it costs several percentage points without showing up as an error.
Code
import numpy as np
from PIL import Image
from torchvision import transforms
IMAGENET = dict(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
train_tf = transforms.Compose([
transforms.RandomResizedCrop(224, scale=(0.7, 1.0)),
transforms.RandomHorizontalFlip(),
transforms.ColorJitter(brightness=0.2, contrast=0.2, saturation=0.2),
transforms.ToTensor(),
transforms.Normalize(**IMAGENET),
])
eval_tf = transforms.Compose([ # deterministic — no randomness
transforms.Resize(256),
transforms.CenterCrop(224),
transforms.ToTensor(),
transforms.Normalize(**IMAGENET),
])
image = Image.open("cat.jpg").convert("RGB") # convert handles greyscale and RGBA
print(np.array(image).shape, np.array(image).dtype) # (512, 640, 3) uint8
x = eval_tf(image)
print(x.shape, round(float(x.mean()), 3)) # torch.Size([3, 224, 224]) -0.021
Augmentations that preserve the label — and those that do not:
| Augmentation | Safe? |
|---|---|
| A horizontal flip | yes, except for text, digits and asymmetric objects |
| Rotation ±15° | yes, except for X-rays and maps where up means something |
| Colour jitter | yes, except when the colour is the class (ripe fruit, warning signs) |
| Cropping | yes, if the object is central — otherwise it can be cut away |
| A vertical flip | rarely; only for satellite images and the like |
The rule: if a human would change the label on the augmented image, the augmentation is wrong.
Mastery means
- Loads, scales and normalises images correctly
- Chooses augmentations that preserve the label
- Avoids the common faults in the pipeline
Sign in to do the exercises and build your mastery up.
Sources
- PyTorch — torchvision transforms (BSD-3) — BSD-3-Clause
- Albumentations (MIT) — MIT