Evaluating dangerous capabilities
Be able to read and criticise safety evals and understand their methodological problems.
Prerequisites
Intuition
Dangerous capabilities are evaluated to answer one question before launch: can this model help somebody bring about serious harm? The areas are in practice: cyber operations, biology and chemistry, large-scale manipulation, and autonomous replication or resource acquisition.
Two things that are often confused:
- Capability: can the model, if somebody really tries?
- Propensity: does it, spontaneously, in normal use?
Safety evals mainly measure the first — and that is the right focus, since guardrails can be circumvented while the capability remains.
Research
The methodological problems, in order:
- The elicitation gap. A negative result means «we could not get the model to do it», not «the model cannot». With better prompts, fine-tuning, tools or scaffolding the performance often rises dramatically. Serious evaluations therefore report which elicitation was used and treat the result as a lower bound.
- Ceiling transparency. If the model scores 60 % on a ceiling-limited test you do not know whether it «really» can do more. The tasks have to have headroom.
- Contamination. Published safety benchmarks end up in the training data just like the others.
- Proxy tasks. For obvious reasons real harm is not tested, but proxy tasks are. How well they predict real capacity is largely untested.
- Situational awareness. A model can behave differently when it believes it is being evaluated — which is documented and makes negative results harder to interpret.
Responsible scaling policies (Anthropic's RSP, OpenAI's Preparedness, Google's Frontier Safety Framework) tie thresholds in such evals to concrete measures: further protection, a limited rollout, or pausing the training. That makes the evaluation methodology a governance question, not just a technical one.
What you can do yourself: read a model card's safety section and ask five questions — which elicitation, which thresholds, who set them, what would have happened at an exceedance, and is there an independent review?
Mastery means
- Reads and criticises a safety eval
- Identifies the methodological problems: elicitation, contamination, ceiling transparency
- Distinguishes capability from propensity
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Model evaluation for extreme risks — arXiv (open access; licence per article)
- Anthropic — Responsible Scaling Policy — free to read
- arXiv — Evaluating Frontier Models for Dangerous Capabilities — arXiv (open access; licence per article)