Skip to content
AI-grafen
FAI engineeringComputer vision· about 90 min· fast-moving, sources checked often· verified 2026-09-21· EN

Object detection

Be able to explain bounding boxes, IoU and NMS, and use a detector.

Prerequisites

Intuition

Classification answers «what is in the image?». Detection answers «what is where?» — a list of boxes with a class and a confidence.

IoU (intersection over union) measures how well two boxes overlap:

IoU=intersection areaunion area\mathrm{IoU} = \frac{\text{intersection area}}{\text{union area}}

IoUMeans
1.0identical boxes
0.5the threshold for a «hit» in the classic metrics
0.0no overlap

NMS (non-maximum suppression) solves the problem that the detector finds the same object several times: sort the boxes by confidence, keep the best one, and discard everything that overlaps it by more than a threshold. Repeat.

Without NMS you get twenty boxes around every car. With NMS you get one.

Formal

Two families:

Two-stage (Faster R-CNN)One-stage (YOLO, RetinaNet, DETR)
Howpropose regions, classify thempredict the boxes directly
Precisionhistorically highernowadays equivalent
Speedslowerreal time
Used forresearch, high precisionin practice nearly always

DETR is a third route: a transformer that predicts a fixed set of boxes and matches them against the ground truth with Hungarian matching. No NMS is needed — the model learns not to duplicate.

mAP (mean average precision) is the standard metric, and it is often misunderstood:

  1. For each class, sort the detections by confidence.
  2. Compute the precision and the recall at every threshold → a precision-recall curve.
  3. AP is the area under that curve.
  4. mAP is the mean over the classes.
  5. mAP@[.5:.95] (the COCO standard) additionally averages over the IoU thresholds 0.50 to 0.95.

The last point is important: [email protected] and mAP@[.5:.95] often differ by 15–20 percentage points, and a comparison that mixes them is meaningless.

Class imbalance is detection's fundamental problem: an image typically has a few objects and tens of thousands of background regions. Two solutions:

SolutionThe idea
Focal lossdown-weight the easy examples so that the hard ones dominate the gradient
Hard negative miningpick out the hardest background examples

Practical choices:

NeedChoice
Real time at the edgeYOLO in a small variant
High precision, not time-criticaltwo-stage or a large DETR variant
Small objectsa higher input resolution — more important than the choice of model
Few labelled imagesa pretrained model plus heavy augmentation

The row about small objects is the one that most often solves the problem in practice: doubling the resolution helps more than swapping the architecture.

Code

import torch, torchvision
from torchvision.ops import box_iou, nms

# IoU by hand
def iou(a, b):
    """a, b: (x1, y1, x2, y2)"""
    x1, y1 = max(a[0], b[0]), max(a[1], b[1])
    x2, y2 = min(a[2], b[2]), min(a[3], b[3])
    intersection = max(0, x2 - x1) * max(0, y2 - y1)
    area_a = (a[2] - a[0]) * (a[3] - a[1])
    area_b = (b[2] - b[0]) * (b[3] - b[1])
    return intersection / (area_a + area_b - intersection)

print(round(iou((0, 0, 10, 10), (5, 5, 15, 15)), 4))     # 0.1429
print(round(iou((0, 0, 10, 10), (0, 0, 10, 10)), 4))     # 1.0
print(round(iou((0, 0, 10, 10), (20, 20, 30, 30)), 4))   # 0.0

# NMS: the same object found several times → keep the best one
boxes = torch.tensor([[10., 10., 50., 50.], [12., 12., 52., 52.],
                      [11., 9., 49., 51.], [100., 100., 140., 140.]])
scores = torch.tensor([0.92, 0.88, 0.85, 0.79])
keep = nms(boxes, scores, iou_threshold=0.5)
print(keep.tolist())            # [0, 3] — three overlapping became one

# Use a pretrained detector
model = torchvision.models.detection.fasterrcnn_resnet50_fpn(weights="DEFAULT").eval()
with torch.no_grad():
    out = model([image])[0]
for box, label, p in zip(out["boxes"], out["labels"], out["scores"]):
    if p > 0.7:
        print(f"  class {int(label)}  confidence {float(p):.2f}  {box.tolist()}")

# mAP: keep [email protected] and mAP@[.5:.95] apart
def ap_at_threshold(pred_boxes, pred_scores, true_boxes, iou_threshold=0.5):
    order = pred_scores.argsort(descending=True)
    pred_boxes = pred_boxes[order]
    matched = torch.zeros(len(true_boxes), dtype=torch.bool)
    tp = torch.zeros(len(pred_boxes))
    for i, p in enumerate(pred_boxes):
        if len(true_boxes) == 0:
            break
        iou_row = box_iou(p[None], true_boxes)[0]
        iou_row[matched] = -1                        # an already used box cannot match again
        best = int(iou_row.argmax())
        if float(iou_row[best]) >= iou_threshold:
            tp[i] = 1.0
            matched[best] = True
    cum_tp = tp.cumsum(0)
    precision = cum_tp / torch.arange(1, len(tp) + 1)
    recall = cum_tp / max(len(true_boxes), 1)
    # 101-point interpolated AP, as in COCO
    ap = 0.0
    for r in torch.linspace(0, 1, 101):
        p_max = precision[recall >= r].max() if (recall >= r).any() else torch.tensor(0.0)
        ap += float(p_max) / 101
    return ap

def map_coco(pred_boxes, pred_scores, true_boxes):
    thresholds = torch.arange(0.5, 1.0, 0.05)
    return {
        "[email protected]": round(ap_at_threshold(pred_boxes, pred_scores, true_boxes, 0.5), 4),
        "mAP@[.5:.95]": round(float(sum(
            ap_at_threshold(pred_boxes, pred_scores, true_boxes, float(t))
            for t in thresholds) / len(thresholds)), 4),
    }

The difference between the two numbers in map_coco is often 15–20 percentage points. Comparing one model's [email protected] with another's mAP@[.5:.95] is one of the most common errors in the detection literature.

Mastery means

  • Explains IoU and NMS
  • Interprets mAP correctly
  • Chooses a detector family according to the requirements

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences