Skip to content
AI-grafen
FAI engineeringAI safety and alignment· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

AI safety and red teaming

Be able to red team a model or an agent, classify the vulnerabilities and propose protections.

Prerequisites

Intuition

Red teaming is systematically trying to make the system do something it should not — before somebody else does.

Four attack surfaces for an LLM system:

SurfaceExamples
Input from the userjailbreaks, role-play, invented authorities, encoding and obfuscation
Input from the worldindirect prompt injection in documents, web pages, email
Tools and side effectsgetting the agent to send, pay, delete, leak
Outputa leaked system prompt, personal data, harmful content

The structure that makes it more than poking about: a checklist of attack classes, a scoring of the findings, and every finding becoming an eval case that runs on every release.

Formal

Classify every finding in two dimensions:

Low likelihoodHigh likelihood
High harmfix before releaseblocks the release
Low harmbacklogfix when the opportunity comes

Likelihood = how easy is it to trigger? An attack that takes an expert 30 attempts is not the same thing as one triggered by an ordinary phrasing.

Protection in layers — no single one is enough:

  1. Least privilege: the model only has the tools the task requires.
  2. Human confirmation for anything irreversible.
  3. Input and output filters: classifiers for known patterns (catching the simple cases).
  4. System prompt hardening: helps marginally, is always circumvented in the end.
  5. Structural separation: retrieved content is marked as data, never as instructions.
  6. Monitoring and a kill switch: detect and shut it down in operation.
  7. An eval suite: every finding becomes a permanent test case.

Ethics and legality: only red team systems you own or have written permission to test. Document the scope and the time window in advance. If you find a vulnerability in somebody else's product: responsible disclosure, not publication.

Research

Automated red teaming scales the manual work: an attacker LLM generates thousands of prompt variants, a classifier judges whether the target model violated the policy, and the successful attacks become training data or eval cases (Perez et al. 2022). Anthropic's and OpenAI's model cards now report such suite results.

Open problems worth knowing about:

  • No known method eliminates prompt injection. It is a structural problem: the model cannot tell instructions from data.
  • Jailbreaks transfer between models — an attack that works on one model often works on others (Zou et al. 2023), including automatically generated suffixes.
  • Capability evaluations (dangerous capabilities) differ from behavioural evaluations and need their own protocols.

For AI-grafen the concrete equivalent is: the tutor must not give away the answer key, must not leave the subject for child accounts, and the lab sandbox must not be escapable — all three are tested in m2-check, isolation_check and the tutor eval suite.

Mastery means

  • Carries out structured red teaming of a model or an agent
  • Classifies the findings by severity and likelihood
  • Proposes protections in several layers

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences