Agent security: authorisations, the sandbox, confirmation
Be able to limit what an agent may do and require confirmation for irreversible actions.
Prerequisites
- EAgents — plan, act, observerequired
- FAI safety and red teamingrequired
Intuition
An agent that can read is an information problem. An agent that can write, send, pay or delete is a security problem.
Classify every tool before it is connected:
| The class | Example | The rule |
|---|---|---|
| Read, public | searching the documentation | free |
| Read, sensitive | customer data, personal data | an authorisation check per user, logged |
| Write, reversible | creating a draft, commenting | allowed, logged, can be undone |
| Write, irreversible | sending an email, paying, deleting | requires a human confirmation of the concrete call |
The confirmation should show exactly what is going to happen — the recipient, the amount, the content — not «the agent wants to send an email, OK?».
Formal
The sandbox's layers (as in AI-grafen's lab runner, ADR-003):
| The layer | The measure | Stops |
|---|---|---|
| Network | --network none | exfiltration, downloading, metadata services |
| File system | a read-only rootfs + tmpfs | persistence, manipulation of the image |
| User | uid 65534, no-new-privileges | privilege escalation |
| Capabilities | cap-drop ALL | kernel operations |
| Resources | cpu, memory, pids, a timeout | resource exhaustion |
| Lifetime | the container is removed after the run | leftover state |
The network isolation is the most important because nearly all damage requires something to leave the machine or be fetched in.
A budget as a security mechanism: max steps, max tokens, max tool calls per run and per user per day. Without it a looping agent can cost more in a night than the service turns over in a month.
Logging for review afterwards: every tool call with its arguments (truncated, masked), the result, a timestamp and who initiated it. Without a trail an incident cannot be investigated.
The boundary that is easy to miss: the agent's authorisation should be the user's authorisation, not the system's. An agent running with the service account's rights can be used by any user at all to reach anything at all.
Code
from dataclasses import dataclass
from enum import Enum
class Risk(Enum):
READ = 1; WRITE_REVERSIBLE = 2; IRREVERSIBLE = 3
@dataclass
class Tool:
name: str
risk: Risk
run: callable
requires_role: str | None = None
async def run_tool(t: Tool, args: dict, user: dict, confirm, log):
if t.requires_role and not has_role(user, t.requires_role):
return {"error": "lacks authorisation"} # the user's authorisation, not the system's
if t.risk is Risk.IRREVERSIBLE:
ok = await confirm(f"{t.name}({summarise(args)})") # show the CONCRETE call
if not ok:
await log("aborted_by_user", t.name, args)
return {"error": "aborted by the user"}
await log("tool_called", t.name, args, user["id"])
return await t.run(args, as_user=user) # run as the user
# A sandbox for running code — the same settings as the platform's lab runner
docker run --rm --network none --read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64m \
--user 65534:65534 --cap-drop ALL --security-opt no-new-privileges \
--cpus 1 --memory 512m --pids-limit 128 \
ai-grafen/lab-python:2026.09 timeout 60 python /work/code.py
Mastery means
- Limits an agent's authorisations to the minimum necessary
- Requires confirmation for irreversible actions
- Runs code in a sandbox with the right restrictions
Sign in to do the exercises and build your mastery up.
Sources
- OWASP Top 10 for LLM Applications — CC BY-SA 4.0
- Docker — security (Apache-2.0) — Apache-2.0