Skip to content
AI-grafen
FAI engineeringComputer vision· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Adversarial examples

Be able to generate adversarial perturbations and explain what they reveal about models.

Prerequisites

Intuition

Take an image classified as «panda» with 58 % confidence. Add a perturbation invisible to the eye — every pixel changed by at most 1/255. Now it is classified as «gibbon» with 99 % confidence.

That is not a fault in that particular model. It is a property of how most networks work: the decision boundaries lie very close to the data points in high-dimensional spaces, and a small move in the right direction is enough.

The direction is found with the gradient — the same gradient used for training, except that it is now used to change the image rather than the weights.

Formal

FGSM (Goodfellow et al. 2014) — a single step: xadv=x+ϵ⋅sign(∇xL(f(x),y))x_{adv} = x + \epsilon\cdot\text{sign}\big(\nabla_x \mathcal L(f(x), y)\big)

PGD — iterative and much stronger: xt+1=Π∥x′−x∥∞≤ϵ(xt+α⋅sign(∇xL(f(xt),y)))x^{t+1} = \Pi_{\|x'-x\|_\infty\le\epsilon}\Big(x^t + \alpha\cdot\text{sign}\big(\nabla_x\mathcal L(f(x^t), y)\big)\Big) The projection Π\Pi keeps the perturbation inside the ϵ\epsilon ball.

What it reveals: Ilyas et al. (2019) argue that adversarial examples exploit non-robust features — patterns that really are predictive in the data but imperceptible to humans. So the model is not «stupid»; it is using signals we do not see.

The defences and their status:

DefenceStatus
Adversarial training (train on PGD examples)works, but costs 3–30× the training time and lowers clean accuracy
Certified defences (randomized smoothing)guarantees, but only for small ε
Gradient masking (obfuscation)does not work — broken by stronger attacks
Input transformationsusually broken by adaptive attacks

The rule in the field: a defence tested only against FGSM has not been tested. Always evaluate against adaptive attacks that know about the defence.

Code

import torch, torch.nn.functional as F

def fgsm(model, x, y, eps=2/255):
    x = x.clone().detach().requires_grad_(True)
    loss = F.cross_entropy(model(x), y)
    loss.backward()
    return (x + eps * x.grad.sign()).clamp(0, 1).detach()

def pgd(model, x, y, eps=2/255, alpha=0.5/255, steps=20):
    x_adv = (x + torch.empty_like(x).uniform_(-eps, eps)).clamp(0, 1).detach()
    for _ in range(steps):
        x_adv.requires_grad_(True)
        loss = F.cross_entropy(model(x_adv), y)
        grad = torch.autograd.grad(loss, x_adv)[0]
        x_adv = x_adv.detach() + alpha * grad.sign()
        x_adv = (x + (x_adv - x).clamp(-eps, eps)).clamp(0, 1)      # project
    return x_adv.detach()

clean = (model(x).argmax(1) == y).float().mean()
adv = (model(pgd(model, x, y)).argmax(1) == y).float().mean()
print(f"clean {clean:.2f} → adversarial {adv:.2f}")     # clean 0.94 → adversarial 0.03

An untrained model's accuracy typically falls from 94 % to near zero at ε = 2/255 — a perturbation that is impossible to see.

Mastery means

  • Generates an adversarial perturbation with FGSM/PGD
  • Explains what the phenomenon reveals about models
  • Knows the defences and their limitations

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences