AI safety and red teaming
Be able to red team a model or an agent, classify the vulnerabilities and propose protections.
Prerequisites
- CResponsible use of AIrequired
- EAgents — plan, act, observerequired
- FEvals for language models and agentsrequired
Intuition
Red teaming is systematically trying to make the system do something it should not — before somebody else does.
Four attack surfaces for an LLM system:
| Surface | Examples |
|---|---|
| Input from the user | jailbreaks, role-play, invented authorities, encoding and obfuscation |
| Input from the world | indirect prompt injection in documents, web pages, email |
| Tools and side effects | getting the agent to send, pay, delete, leak |
| Output | a leaked system prompt, personal data, harmful content |
The structure that makes it more than poking about: a checklist of attack classes, a scoring of the findings, and every finding becoming an eval case that runs on every release.
Formal
Classify every finding in two dimensions:
| Low likelihood | High likelihood | |
|---|---|---|
| High harm | fix before release | blocks the release |
| Low harm | backlog | fix when the opportunity comes |
Likelihood = how easy is it to trigger? An attack that takes an expert 30 attempts is not the same thing as one triggered by an ordinary phrasing.
Protection in layers — no single one is enough:
- Least privilege: the model only has the tools the task requires.
- Human confirmation for anything irreversible.
- Input and output filters: classifiers for known patterns (catching the simple cases).
- System prompt hardening: helps marginally, is always circumvented in the end.
- Structural separation: retrieved content is marked as data, never as instructions.
- Monitoring and a kill switch: detect and shut it down in operation.
- An eval suite: every finding becomes a permanent test case.
Ethics and legality: only red team systems you own or have written permission to test. Document the scope and the time window in advance. If you find a vulnerability in somebody else's product: responsible disclosure, not publication.
Research
Automated red teaming scales the manual work: an attacker LLM generates thousands of prompt variants, a classifier judges whether the target model violated the policy, and the successful attacks become training data or eval cases (Perez et al. 2022). Anthropic's and OpenAI's model cards now report such suite results.
Open problems worth knowing about:
- No known method eliminates prompt injection. It is a structural problem: the model cannot tell instructions from data.
- Jailbreaks transfer between models — an attack that works on one model often works on others (Zou et al. 2023), including automatically generated suffixes.
- Capability evaluations (dangerous capabilities) differ from behavioural evaluations and need their own protocols.
For AI-grafen the concrete equivalent is: the tutor must not give away the answer key, must not leave the subject for child accounts, and the lab sandbox must not be escapable — all three are tested in m2-check, isolation_check and the tutor eval suite.
Mastery means
- Carries out structured red teaming of a model or an agent
- Classifies the findings by severity and likelihood
- Proposes protections in several layers
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Red Teaming Language Models with Language Models — arXiv (open access; licence per article)
- arXiv — Universal and Transferable Adversarial Attacks on Aligned Language Models — arXiv (open access; licence per article)
- OWASP Top 10 for LLM Applications — CC BY-SA 4.0