Skip to content
AI-grafen
FAI engineeringAI safety and alignment· about 90 min· fast-moving, sources checked often· verified 2026-09-21· EN

Data poisoning and backdoors

Be able to explain how training data can be manipulated and how it is detected.

Prerequisites

Intuition

Data poisoning is manipulating the training data so that the model learns something the attacker wants.

AttackThe goalVisible in the evaluation?
Availability attackdegrade the model generallyyes
Targeted attackproduce wrong answers on specific inputsbarely
Backdoornormal behaviour, but wrong on a triggerno

The backdoor is the unpleasant one. The model behaves entirely normally — except when a particular pattern appears in the input, and then it does something else.

A classic example: an image classifier where a small yellow square in the corner makes everything get classified as «speed limit 80». On all ordinary images the model works perfectly. The test set shows nothing.

Why it is relevant now: modern models are trained on web-scraped data that anybody can contribute to. Carlini et al. (2023) showed that it is practically feasible to poison a small but sufficient share of several well-known web corpora — among other things by buying expired domains that are part of them.

Formal

The attack surface in a modern pipeline:

SurfaceHow
The web corpuspublish text that gets scraped; buy expired domains
Crowdsourced labellingmislabel systematically
User-generated dataif the system trains on its own traffic
RLHF preferencesmanipulate the rankings
Pretrained weightsa model from an unknown source can contain a backdoor
Dependenciesa library somewhere in the chain

The second to last is the most underestimated: downloading a model from a public hub is trusting whoever uploaded it. Check the signatures and use safetensors — pickle-based formats can in addition execute code on load, which is a different and more immediate problem.

How little it takes. The research points to a surprisingly small share of poisoned examples being enough for a targeted backdoor — often fractions of a per cent, and in some set-ups a fixed number of examples rather than a share. That means «we have billions of documents, a few bad ones make no difference» is not a valid argument.

The countermeasures, and what they actually manage:

CountermeasureManagesDoes not manage
Provenance and signingunknown sourcesa poisoned but signed source
Deduplicationrepeated injectionssingle examples
Activation clusteringbackdoors with a clear signaturesubtle triggers
Spectral signature analysisoutliers in the representation spacesmall shares
Fine-tuning on clean datamany backdoorspersistent ones
Pruning unused neuronssome backdoors
Trigger searchknown kinds of patternunknown ones

The honest summary is that no method gives guarantees. The defence rests on layers: known provenance, filtering, anomaly detection and evaluation that actively looks for aberrant behaviour.

What you do in practice:

  1. Know where the data comes from. Log the source per document.
  2. Sign and verify the model weights and the datasets.
  3. Do not train on your own production traffic without review.
  4. Evaluate on held-out, trusted test sets that have never been in contact with the pipeline.
  5. Actively look for backdoors in third-party models before they go into production.

Point 3 is particularly relevant for systems that collect user interactions: a feedback loop that trains on what users submit is an open invitation.

Code

import numpy as np, torch
from collections import Counter

# Activation clustering: poisoned examples often form a cluster of their own
def activation_clustering(model, dataloader, layer, n_clusters=2):
    from sklearn.cluster import KMeans
    from sklearn.decomposition import PCA
    stored = []
    h = layer.register_forward_hook(lambda m, i, o: stored.append(o.detach().flatten(1).cpu()))
    labels = []
    with torch.no_grad():
        for x, y in dataloader:
            model(x); labels.append(y)
    h.remove()
    A = torch.cat(stored).numpy()
    y = torch.cat(labels).numpy()

    suspicious = {}
    for cls in np.unique(y):
        mask = y == cls
        if mask.sum() < 20:
            continue
        Z = PCA(n_components=10).fit_transform(A[mask])
        k = KMeans(n_clusters=n_clusters, n_init=10, random_state=0).fit(Z)
        shares = np.bincount(k.labels_, minlength=n_clusters) / mask.sum()
        if shares.min() < 0.15:           # a small, clearly separated cluster
            suspicious[int(cls)] = round(float(shares.min()), 4)
    return suspicious
# {3: 0.042}  ← 4 % of class 3 sits in a cluster of its own: investigate them

# A trigger test: measure whether a pattern systematically changes the classification
def trigger_test(model, images, trigger_fn, n=200):
    model.eval()
    with torch.no_grad():
        before = model(images[:n]).argmax(1)
        after = model(torch.stack([trigger_fn(b) for b in images[:n]])).argmax(1)
    changed = (before != after)
    targets = Counter(int(k) for k in after[changed])
    dominant = targets.most_common(1)[0] if targets else (None, 0)
    return {"share_changed": round(float(changed.float().mean()), 3),
            "dominant_target": dominant[0],
            "share_to_that_target": round(dominant[1] / max(int(changed.sum()), 1), 3)}
# A high share changed AND a dominant target class = a strong indication of a backdoor

def yellow_square(image, size=4):
    b = image.clone()
    b[:, -size:, -size:] = torch.tensor([1.0, 1.0, 0.0])[:, None, None]
    return b

# Provenance: log the source per document and make it auditable
def corpus_with_provenance(documents):
    import hashlib
    return [{"text": d["text"],
             "source": d["source"],
             "fetched": d["fetched"],
             "sha256": hashlib.sha256(d["text"].encode()).hexdigest()[:16],
             "trusted": d["source"] in TRUSTED_SOURCES}
            for d in documents]

def untrusted_share(corpus):
    n = sum(1 for d in corpus if not d["trusted"])
    return {"untrusted": n, "of": len(corpus), "share": round(n / max(len(corpus), 1), 4)}

# Load models safely
from safetensors.torch import load_file
weights = load_file("model.safetensors")     # cannot by construction execute code
# torch.load("model.pt")  ← runs pickle; never from an unknown source

Mastery means

  • Distinguishes the different kinds of attack on training data
  • Explains how backdoors work
  • Describes realistic countermeasures

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences