Object detection
Be able to explain bounding boxes, IoU and NMS, and use a detector.
Prerequisites
- ECNN architectures: LeNet to ResNetrequired
Intuition
Classification answers «what is in the image?». Detection answers «what is where?» — a list of boxes with a class and a confidence.
IoU (intersection over union) measures how well two boxes overlap:
| IoU | Means |
|---|---|
| 1.0 | identical boxes |
| 0.5 | the threshold for a «hit» in the classic metrics |
| 0.0 | no overlap |
NMS (non-maximum suppression) solves the problem that the detector finds the same object several times: sort the boxes by confidence, keep the best one, and discard everything that overlaps it by more than a threshold. Repeat.
Without NMS you get twenty boxes around every car. With NMS you get one.
Formal
Two families:
| Two-stage (Faster R-CNN) | One-stage (YOLO, RetinaNet, DETR) | |
|---|---|---|
| How | propose regions, classify them | predict the boxes directly |
| Precision | historically higher | nowadays equivalent |
| Speed | slower | real time |
| Used for | research, high precision | in practice nearly always |
DETR is a third route: a transformer that predicts a fixed set of boxes and matches them against the ground truth with Hungarian matching. No NMS is needed — the model learns not to duplicate.
mAP (mean average precision) is the standard metric, and it is often misunderstood:
- For each class, sort the detections by confidence.
- Compute the precision and the recall at every threshold → a precision-recall curve.
- AP is the area under that curve.
- mAP is the mean over the classes.
- mAP@[.5:.95] (the COCO standard) additionally averages over the IoU thresholds 0.50 to 0.95.
The last point is important: [email protected] and mAP@[.5:.95] often differ by 15–20 percentage points, and a comparison that mixes them is meaningless.
Class imbalance is detection's fundamental problem: an image typically has a few objects and tens of thousands of background regions. Two solutions:
| Solution | The idea |
|---|---|
| Focal loss | down-weight the easy examples so that the hard ones dominate the gradient |
| Hard negative mining | pick out the hardest background examples |
Practical choices:
| Need | Choice |
|---|---|
| Real time at the edge | YOLO in a small variant |
| High precision, not time-critical | two-stage or a large DETR variant |
| Small objects | a higher input resolution — more important than the choice of model |
| Few labelled images | a pretrained model plus heavy augmentation |
The row about small objects is the one that most often solves the problem in practice: doubling the resolution helps more than swapping the architecture.
Code
import torch, torchvision
from torchvision.ops import box_iou, nms
# IoU by hand
def iou(a, b):
"""a, b: (x1, y1, x2, y2)"""
x1, y1 = max(a[0], b[0]), max(a[1], b[1])
x2, y2 = min(a[2], b[2]), min(a[3], b[3])
intersection = max(0, x2 - x1) * max(0, y2 - y1)
area_a = (a[2] - a[0]) * (a[3] - a[1])
area_b = (b[2] - b[0]) * (b[3] - b[1])
return intersection / (area_a + area_b - intersection)
print(round(iou((0, 0, 10, 10), (5, 5, 15, 15)), 4)) # 0.1429
print(round(iou((0, 0, 10, 10), (0, 0, 10, 10)), 4)) # 1.0
print(round(iou((0, 0, 10, 10), (20, 20, 30, 30)), 4)) # 0.0
# NMS: the same object found several times → keep the best one
boxes = torch.tensor([[10., 10., 50., 50.], [12., 12., 52., 52.],
[11., 9., 49., 51.], [100., 100., 140., 140.]])
scores = torch.tensor([0.92, 0.88, 0.85, 0.79])
keep = nms(boxes, scores, iou_threshold=0.5)
print(keep.tolist()) # [0, 3] — three overlapping became one
# Use a pretrained detector
model = torchvision.models.detection.fasterrcnn_resnet50_fpn(weights="DEFAULT").eval()
with torch.no_grad():
out = model([image])[0]
for box, label, p in zip(out["boxes"], out["labels"], out["scores"]):
if p > 0.7:
print(f" class {int(label)} confidence {float(p):.2f} {box.tolist()}")
# mAP: keep [email protected] and mAP@[.5:.95] apart
def ap_at_threshold(pred_boxes, pred_scores, true_boxes, iou_threshold=0.5):
order = pred_scores.argsort(descending=True)
pred_boxes = pred_boxes[order]
matched = torch.zeros(len(true_boxes), dtype=torch.bool)
tp = torch.zeros(len(pred_boxes))
for i, p in enumerate(pred_boxes):
if len(true_boxes) == 0:
break
iou_row = box_iou(p[None], true_boxes)[0]
iou_row[matched] = -1 # an already used box cannot match again
best = int(iou_row.argmax())
if float(iou_row[best]) >= iou_threshold:
tp[i] = 1.0
matched[best] = True
cum_tp = tp.cumsum(0)
precision = cum_tp / torch.arange(1, len(tp) + 1)
recall = cum_tp / max(len(true_boxes), 1)
# 101-point interpolated AP, as in COCO
ap = 0.0
for r in torch.linspace(0, 1, 101):
p_max = precision[recall >= r].max() if (recall >= r).any() else torch.tensor(0.0)
ap += float(p_max) / 101
return ap
def map_coco(pred_boxes, pred_scores, true_boxes):
thresholds = torch.arange(0.5, 1.0, 0.05)
return {
"[email protected]": round(ap_at_threshold(pred_boxes, pred_scores, true_boxes, 0.5), 4),
"mAP@[.5:.95]": round(float(sum(
ap_at_threshold(pred_boxes, pred_scores, true_boxes, float(t))
for t in thresholds) / len(thresholds)), 4),
}
The difference between the two numbers in map_coco is often 15–20 percentage points. Comparing one model's [email protected] with another's mAP@[.5:.95] is one of the most common errors in the detection literature.
Mastery means
- Explains IoU and NMS
- Interprets mAP correctly
- Chooses a detector family according to the requirements
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Focal Loss for Dense Object Detection — arXiv (open access; licence per article)
- arXiv — End-to-End Object Detection with Transformers (DETR) — arXiv (open access; licence per article)
- PyTorch — tutorials (BSD-3) — BSD-3-Clause