Adversarial examples
Be able to generate adversarial perturbations and explain what they reveal about models.
Prerequisites
Intuition
Take an image classified as «panda» with 58 % confidence. Add a perturbation invisible to the eye — every pixel changed by at most 1/255. Now it is classified as «gibbon» with 99 % confidence.
That is not a fault in that particular model. It is a property of how most networks work: the decision boundaries lie very close to the data points in high-dimensional spaces, and a small move in the right direction is enough.
The direction is found with the gradient — the same gradient used for training, except that it is now used to change the image rather than the weights.
Formal
FGSM (Goodfellow et al. 2014) — a single step:
PGD — iterative and much stronger: The projection keeps the perturbation inside the ball.
What it reveals: Ilyas et al. (2019) argue that adversarial examples exploit non-robust features — patterns that really are predictive in the data but imperceptible to humans. So the model is not «stupid»; it is using signals we do not see.
The defences and their status:
| Defence | Status |
|---|---|
| Adversarial training (train on PGD examples) | works, but costs 3–30× the training time and lowers clean accuracy |
| Certified defences (randomized smoothing) | guarantees, but only for small ε |
| Gradient masking (obfuscation) | does not work — broken by stronger attacks |
| Input transformations | usually broken by adaptive attacks |
The rule in the field: a defence tested only against FGSM has not been tested. Always evaluate against adaptive attacks that know about the defence.
Code
import torch, torch.nn.functional as F
def fgsm(model, x, y, eps=2/255):
x = x.clone().detach().requires_grad_(True)
loss = F.cross_entropy(model(x), y)
loss.backward()
return (x + eps * x.grad.sign()).clamp(0, 1).detach()
def pgd(model, x, y, eps=2/255, alpha=0.5/255, steps=20):
x_adv = (x + torch.empty_like(x).uniform_(-eps, eps)).clamp(0, 1).detach()
for _ in range(steps):
x_adv.requires_grad_(True)
loss = F.cross_entropy(model(x_adv), y)
grad = torch.autograd.grad(loss, x_adv)[0]
x_adv = x_adv.detach() + alpha * grad.sign()
x_adv = (x + (x_adv - x).clamp(-eps, eps)).clamp(0, 1) # project
return x_adv.detach()
clean = (model(x).argmax(1) == y).float().mean()
adv = (model(pgd(model, x, y)).argmax(1) == y).float().mean()
print(f"clean {clean:.2f} → adversarial {adv:.2f}") # clean 0.94 → adversarial 0.03
An untrained model's accuracy typically falls from 94 % to near zero at ε = 2/255 — a perturbation that is impossible to see.
Mastery means
- Generates an adversarial perturbation with FGSM/PGD
- Explains what the phenomenon reveals about models
- Knows the defences and their limitations
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Explaining and Harnessing Adversarial Examples — arXiv (open access; licence per article)
- arXiv — Towards Deep Learning Models Resistant to Adversarial Attacks (PGD) — arXiv (open access; licence per article)
- arXiv — Adversarial Examples Are Not Bugs, They Are Features — arXiv (open access; licence per article)