Skip to content
AI-grafen
FAI engineeringAgents and tool use· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Agent security: authorisations, the sandbox, confirmation

Be able to limit what an agent may do and require confirmation for irreversible actions.

Prerequisites

Intuition

An agent that can read is an information problem. An agent that can write, send, pay or delete is a security problem.

Classify every tool before it is connected:

The classExampleThe rule
Read, publicsearching the documentationfree
Read, sensitivecustomer data, personal dataan authorisation check per user, logged
Write, reversiblecreating a draft, commentingallowed, logged, can be undone
Write, irreversiblesending an email, paying, deletingrequires a human confirmation of the concrete call

The confirmation should show exactly what is going to happen — the recipient, the amount, the content — not «the agent wants to send an email, OK?».

Formal

The sandbox's layers (as in AI-grafen's lab runner, ADR-003):

The layerThe measureStops
Network--network noneexfiltration, downloading, metadata services
File systema read-only rootfs + tmpfspersistence, manipulation of the image
Useruid 65534, no-new-privilegesprivilege escalation
Capabilitiescap-drop ALLkernel operations
Resourcescpu, memory, pids, a timeoutresource exhaustion
Lifetimethe container is removed after the runleftover state

The network isolation is the most important because nearly all damage requires something to leave the machine or be fetched in.

A budget as a security mechanism: max steps, max tokens, max tool calls per run and per user per day. Without it a looping agent can cost more in a night than the service turns over in a month.

Logging for review afterwards: every tool call with its arguments (truncated, masked), the result, a timestamp and who initiated it. Without a trail an incident cannot be investigated.

The boundary that is easy to miss: the agent's authorisation should be the user's authorisation, not the system's. An agent running with the service account's rights can be used by any user at all to reach anything at all.

Code

from dataclasses import dataclass
from enum import Enum

class Risk(Enum):
    READ = 1; WRITE_REVERSIBLE = 2; IRREVERSIBLE = 3

@dataclass
class Tool:
    name: str
    risk: Risk
    run: callable
    requires_role: str | None = None

async def run_tool(t: Tool, args: dict, user: dict, confirm, log):
    if t.requires_role and not has_role(user, t.requires_role):
        return {"error": "lacks authorisation"}                   # the user's authorisation, not the system's
    if t.risk is Risk.IRREVERSIBLE:
        ok = await confirm(f"{t.name}({summarise(args)})")        # show the CONCRETE call
        if not ok:
            await log("aborted_by_user", t.name, args)
            return {"error": "aborted by the user"}
    await log("tool_called", t.name, args, user["id"])
    return await t.run(args, as_user=user)                        # run as the user
# A sandbox for running code — the same settings as the platform's lab runner
docker run --rm --network none --read-only \
  --tmpfs /tmp:rw,noexec,nosuid,size=64m \
  --user 65534:65534 --cap-drop ALL --security-opt no-new-privileges \
  --cpus 1 --memory 512m --pids-limit 128 \
  ai-grafen/lab-python:2026.09 timeout 60 python /work/code.py

Mastery means

  • Limits an agent's authorisations to the minimum necessary
  • Requires confirmation for irreversible actions
  • Runs code in a sandbox with the right restrictions

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences