Skip to content
AI-grafen
GFrontier LabAI safety and alignment· about 120 min· fast-moving, sources checked often· verified 2026-09-20· EN

Evaluating dangerous capabilities

Be able to read and criticise safety evals and understand their methodological problems.

Prerequisites

Intuition

Dangerous capabilities are evaluated to answer one question before launch: can this model help somebody bring about serious harm? The areas are in practice: cyber operations, biology and chemistry, large-scale manipulation, and autonomous replication or resource acquisition.

Two things that are often confused:

  • Capability: can the model, if somebody really tries?
  • Propensity: does it, spontaneously, in normal use?

Safety evals mainly measure the first — and that is the right focus, since guardrails can be circumvented while the capability remains.

Research

The methodological problems, in order:

  1. The elicitation gap. A negative result means «we could not get the model to do it», not «the model cannot». With better prompts, fine-tuning, tools or scaffolding the performance often rises dramatically. Serious evaluations therefore report which elicitation was used and treat the result as a lower bound.
  2. Ceiling transparency. If the model scores 60 % on a ceiling-limited test you do not know whether it «really» can do more. The tasks have to have headroom.
  3. Contamination. Published safety benchmarks end up in the training data just like the others.
  4. Proxy tasks. For obvious reasons real harm is not tested, but proxy tasks are. How well they predict real capacity is largely untested.
  5. Situational awareness. A model can behave differently when it believes it is being evaluated — which is documented and makes negative results harder to interpret.

Responsible scaling policies (Anthropic's RSP, OpenAI's Preparedness, Google's Frontier Safety Framework) tie thresholds in such evals to concrete measures: further protection, a limited rollout, or pausing the training. That makes the evaluation methodology a governance question, not just a technical one.

What you can do yourself: read a model card's safety section and ask five questions — which elicitation, which thresholds, who set them, what would have happened at an exceedance, and is there an independent review?

Mastery means

  • Reads and criticises a safety eval
  • Identifies the methodological problems: elicitation, contamination, ceiling transparency
  • Distinguishes capability from propensity

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences