Data poisoning and backdoors
Be able to explain how training data can be manipulated and how it is detected.
Prerequisites
Intuition
Data poisoning is manipulating the training data so that the model learns something the attacker wants.
| Attack | The goal | Visible in the evaluation? |
|---|---|---|
| Availability attack | degrade the model generally | yes |
| Targeted attack | produce wrong answers on specific inputs | barely |
| Backdoor | normal behaviour, but wrong on a trigger | no |
The backdoor is the unpleasant one. The model behaves entirely normally — except when a particular pattern appears in the input, and then it does something else.
A classic example: an image classifier where a small yellow square in the corner makes everything get classified as «speed limit 80». On all ordinary images the model works perfectly. The test set shows nothing.
Why it is relevant now: modern models are trained on web-scraped data that anybody can contribute to. Carlini et al. (2023) showed that it is practically feasible to poison a small but sufficient share of several well-known web corpora — among other things by buying expired domains that are part of them.
Formal
The attack surface in a modern pipeline:
| Surface | How |
|---|---|
| The web corpus | publish text that gets scraped; buy expired domains |
| Crowdsourced labelling | mislabel systematically |
| User-generated data | if the system trains on its own traffic |
| RLHF preferences | manipulate the rankings |
| Pretrained weights | a model from an unknown source can contain a backdoor |
| Dependencies | a library somewhere in the chain |
The second to last is the most underestimated: downloading a model from a public hub is trusting whoever uploaded it. Check the signatures and use safetensors — pickle-based formats can in addition execute code on load, which is a different and more immediate problem.
How little it takes. The research points to a surprisingly small share of poisoned examples being enough for a targeted backdoor — often fractions of a per cent, and in some set-ups a fixed number of examples rather than a share. That means «we have billions of documents, a few bad ones make no difference» is not a valid argument.
The countermeasures, and what they actually manage:
| Countermeasure | Manages | Does not manage |
|---|---|---|
| Provenance and signing | unknown sources | a poisoned but signed source |
| Deduplication | repeated injections | single examples |
| Activation clustering | backdoors with a clear signature | subtle triggers |
| Spectral signature analysis | outliers in the representation space | small shares |
| Fine-tuning on clean data | many backdoors | persistent ones |
| Pruning unused neurons | some backdoors | |
| Trigger search | known kinds of pattern | unknown ones |
The honest summary is that no method gives guarantees. The defence rests on layers: known provenance, filtering, anomaly detection and evaluation that actively looks for aberrant behaviour.
What you do in practice:
- Know where the data comes from. Log the source per document.
- Sign and verify the model weights and the datasets.
- Do not train on your own production traffic without review.
- Evaluate on held-out, trusted test sets that have never been in contact with the pipeline.
- Actively look for backdoors in third-party models before they go into production.
Point 3 is particularly relevant for systems that collect user interactions: a feedback loop that trains on what users submit is an open invitation.
Code
import numpy as np, torch
from collections import Counter
# Activation clustering: poisoned examples often form a cluster of their own
def activation_clustering(model, dataloader, layer, n_clusters=2):
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
stored = []
h = layer.register_forward_hook(lambda m, i, o: stored.append(o.detach().flatten(1).cpu()))
labels = []
with torch.no_grad():
for x, y in dataloader:
model(x); labels.append(y)
h.remove()
A = torch.cat(stored).numpy()
y = torch.cat(labels).numpy()
suspicious = {}
for cls in np.unique(y):
mask = y == cls
if mask.sum() < 20:
continue
Z = PCA(n_components=10).fit_transform(A[mask])
k = KMeans(n_clusters=n_clusters, n_init=10, random_state=0).fit(Z)
shares = np.bincount(k.labels_, minlength=n_clusters) / mask.sum()
if shares.min() < 0.15: # a small, clearly separated cluster
suspicious[int(cls)] = round(float(shares.min()), 4)
return suspicious
# {3: 0.042} ← 4 % of class 3 sits in a cluster of its own: investigate them
# A trigger test: measure whether a pattern systematically changes the classification
def trigger_test(model, images, trigger_fn, n=200):
model.eval()
with torch.no_grad():
before = model(images[:n]).argmax(1)
after = model(torch.stack([trigger_fn(b) for b in images[:n]])).argmax(1)
changed = (before != after)
targets = Counter(int(k) for k in after[changed])
dominant = targets.most_common(1)[0] if targets else (None, 0)
return {"share_changed": round(float(changed.float().mean()), 3),
"dominant_target": dominant[0],
"share_to_that_target": round(dominant[1] / max(int(changed.sum()), 1), 3)}
# A high share changed AND a dominant target class = a strong indication of a backdoor
def yellow_square(image, size=4):
b = image.clone()
b[:, -size:, -size:] = torch.tensor([1.0, 1.0, 0.0])[:, None, None]
return b
# Provenance: log the source per document and make it auditable
def corpus_with_provenance(documents):
import hashlib
return [{"text": d["text"],
"source": d["source"],
"fetched": d["fetched"],
"sha256": hashlib.sha256(d["text"].encode()).hexdigest()[:16],
"trusted": d["source"] in TRUSTED_SOURCES}
for d in documents]
def untrusted_share(corpus):
n = sum(1 for d in corpus if not d["trusted"])
return {"untrusted": n, "of": len(corpus), "share": round(n / max(len(corpus), 1), 4)}
# Load models safely
from safetensors.torch import load_file
weights = load_file("model.safetensors") # cannot by construction execute code
# torch.load("model.pt") ← runs pickle; never from an unknown source
Mastery means
- Distinguishes the different kinds of attack on training data
- Explains how backdoors work
- Describes realistic countermeasures
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Poisoning Web-Scale Training Datasets is Practical — arXiv (open access; licence per article)
- arXiv — BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain — arXiv (open access; licence per article)
- OWASP Top 10 for LLM Applications — CC BY-SA 4.0