Skip to content
AI-grafen
EUniversityComputer vision· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

Computer vision — the basics

Be able to train an image classifier, use data augmentation and explain transfer learning.

Prerequisites

Intuition

The recipe for image classification in practice — almost never from scratch:

  1. A pretrained model (ResNet, EfficientNet, ViT) as the base.
  2. Augmentation on the training data: random cropping, flipping, colour jitter, rotation. Never on validation and test.
  3. Normalisation with the pretraining statistics.
  4. A new head for your classes; freeze or fine-tune depending on how much data you have.
  5. Evaluate per class with a confusion matrix, not just the total accuracy.
  6. Error analysis: look at 30 misclassified images. It is always instructive and usually leads to a fix in the data.

Augmentation is the cheapest way to get «more data» — and often what gives the biggest lift on small datasets.

Code

import torch, torch.nn as nn
from torchvision import models, transforms, datasets

IMAGENET = dict(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])

train_tf = transforms.Compose([
    transforms.RandomResizedCrop(224, scale=(0.7, 1.0)),
    transforms.RandomHorizontalFlip(),
    transforms.ColorJitter(0.2, 0.2, 0.2),
    transforms.ToTensor(), transforms.Normalize(**IMAGENET),
])
eval_tf = transforms.Compose([                      # no randomness at evaluation
    transforms.Resize(256), transforms.CenterCrop(224),
    transforms.ToTensor(), transforms.Normalize(**IMAGENET),
])

m = models.resnet18(weights="IMAGENET1K_V1")
for p in m.parameters():
    p.requires_grad = False
m.fc = nn.Linear(m.fc.in_features, num_classes)

# … train the head, then measure per class:
from sklearn.metrics import classification_report, confusion_matrix
print(classification_report(y_true, y_pred, target_names=class_names))
print(confusion_matrix(y_true, y_pred))

Augmentation that hurts: horizontal flips on text or digits (a mirrored «3» is not a «3»), heavy colour jitter when the colour is the class (ripe/unripe fruit), and rotation on images where up means something (X-rays). Only augment with transformations that preserve the label.

Mastery means

  • Trains an image classifier with transfer learning
  • Uses data augmentation correctly
  • Evaluates per class and analyses the mistakes

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences