Skip to content
AI-grafen
EUniversityData handling· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Image preprocessing and augmentation

Be able to load, scale, normalise and augment images.

Prerequisites

Intuition

Before an image reaches the model it goes through four steps, in this order:

  1. Load and decode (JPEG → an array). Check the channel order: OpenCV gives BGR, PIL gives RGB. That is a classic and silent bug.
  2. Scale it to the size the model expects. Preserve the proportions if the shape means something — otherwise objects come out stretched.
  3. Normalise with the same mean and standard deviation the pretraining used.
  4. Augment — but only on the training data.

Step 3 is the one most often forgotten in transfer learning, and it costs several percentage points without showing up as an error.

Code

import numpy as np
from PIL import Image
from torchvision import transforms

IMAGENET = dict(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])

train_tf = transforms.Compose([
    transforms.RandomResizedCrop(224, scale=(0.7, 1.0)),
    transforms.RandomHorizontalFlip(),
    transforms.ColorJitter(brightness=0.2, contrast=0.2, saturation=0.2),
    transforms.ToTensor(),
    transforms.Normalize(**IMAGENET),
])

eval_tf = transforms.Compose([                 # deterministic — no randomness
    transforms.Resize(256),
    transforms.CenterCrop(224),
    transforms.ToTensor(),
    transforms.Normalize(**IMAGENET),
])

image = Image.open("cat.jpg").convert("RGB")   # convert handles greyscale and RGBA
print(np.array(image).shape, np.array(image).dtype)    # (512, 640, 3) uint8
x = eval_tf(image)
print(x.shape, round(float(x.mean()), 3))              # torch.Size([3, 224, 224]) -0.021

Augmentations that preserve the label — and those that do not:

AugmentationSafe?
A horizontal flipyes, except for text, digits and asymmetric objects
Rotation ±15°yes, except for X-rays and maps where up means something
Colour jitteryes, except when the colour is the class (ripe fruit, warning signs)
Croppingyes, if the object is central — otherwise it can be cut away
A vertical fliprarely; only for satellite images and the like

The rule: if a human would change the label on the augmented image, the augmentation is wrong.

Mastery means

  • Loads, scales and normalises images correctly
  • Chooses augmentations that preserve the label
  • Avoids the common faults in the pipeline

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences