Computer vision — the basics
Be able to train an image classifier, use data augmentation and explain transfer learning.
Prerequisites
- DTraining, validation and testrequired
- EConvolutional networks (CNNs)required
Intuition
The recipe for image classification in practice — almost never from scratch:
- A pretrained model (ResNet, EfficientNet, ViT) as the base.
- Augmentation on the training data: random cropping, flipping, colour jitter, rotation. Never on validation and test.
- Normalisation with the pretraining statistics.
- A new head for your classes; freeze or fine-tune depending on how much data you have.
- Evaluate per class with a confusion matrix, not just the total accuracy.
- Error analysis: look at 30 misclassified images. It is always instructive and usually leads to a fix in the data.
Augmentation is the cheapest way to get «more data» — and often what gives the biggest lift on small datasets.
Code
import torch, torch.nn as nn
from torchvision import models, transforms, datasets
IMAGENET = dict(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
train_tf = transforms.Compose([
transforms.RandomResizedCrop(224, scale=(0.7, 1.0)),
transforms.RandomHorizontalFlip(),
transforms.ColorJitter(0.2, 0.2, 0.2),
transforms.ToTensor(), transforms.Normalize(**IMAGENET),
])
eval_tf = transforms.Compose([ # no randomness at evaluation
transforms.Resize(256), transforms.CenterCrop(224),
transforms.ToTensor(), transforms.Normalize(**IMAGENET),
])
m = models.resnet18(weights="IMAGENET1K_V1")
for p in m.parameters():
p.requires_grad = False
m.fc = nn.Linear(m.fc.in_features, num_classes)
# … train the head, then measure per class:
from sklearn.metrics import classification_report, confusion_matrix
print(classification_report(y_true, y_pred, target_names=class_names))
print(confusion_matrix(y_true, y_pred))
Augmentation that hurts: horizontal flips on text or digits (a mirrored «3» is not a «3»), heavy colour jitter when the colour is the class (ripe/unripe fruit), and rotation on images where up means something (X-rays). Only augment with transformations that preserve the label.
Mastery means
- Trains an image classifier with transfer learning
- Uses data augmentation correctly
- Evaluates per class and analyses the mistakes
Sign in to do the exercises and build your mastery up.
Sources
- PyTorch — torchvision (BSD-3) — BSD-3-Clause
- Dive into Deep Learning (CC BY-SA 4.0) — CC BY-SA 4.0