Skip to content
AI-grafen
FAI engineeringAgents and tool use· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Evaluating agents

Be able to build task-based evals for agents with success criteria and cost.

Prerequisites

Intuition

Agents cannot be evaluated like chat answers. The question is not «did the answer sound good?» but «was the task solved?».

Build evals with a verifiable end state:

The taskThe verification
«book the meeting»is the entry in the calendar with the right time?
«fix the bug»does the test suite pass?
«find the price»does the answer match the key?
«clean the data file»does the file satisfy the schema?

Run in a sandbox with a controlled starting state, let the agent work, inspect the end state with code. No judge is needed — it is binary.

Always report three numbers together: the solution rate, the cost per successful task, and the share aborted.

Code

from dataclasses import dataclass
from typing import Callable

@dataclass
class AgentCase:
    id: str
    task: str
    start_state: Callable[[], object]        # builds the sandbox
    verify: Callable[[object], bool]         # inspects the end state

def run_eval_suite(agent, cases: list[AgentCase], attempts: int = 3):
    rows = []
    for c in cases:
        for k in range(attempts):                     # several attempts: agents are stochastic
            world = c.start_state()
            res = agent.run(c.task, world)
            rows.append({"case": c.id, "attempt": k, "succeeded": bool(c.verify(world)),
                         "steps": res["steps"], "tokens": res["tokens"], "sec": res["sec"],
                         "status": res["status"]})
    n = len(rows); ok = [r for r in rows if r["succeeded"]]
    return {
        "solution_rate": len(ok) / n,
        "pass_at_1": sum(r["succeeded"] for r in rows if r["attempt"] == 0) / len(cases),
        "cost_per_success": sum(r["tokens"] for r in rows) / max(len(ok), 1),
        "mean_steps_successful": sum(r["steps"] for r in ok) / max(len(ok), 1),
        "aborted": sum(r["status"] != "done" for r in rows) / n,
    }

Two traps:

  • Measuring the process. «Did the agent use the right tools in the right order?» sounds reasonable but punishes smarter solutions. Measure the end state.
  • One attempt per case. Agents are stochastic; pass@1 and pass@3 say different things and both are relevant.

Mastery means

  • Builds task-based evals with a verifiable end state
  • Measures the cost and the steps, not just success
  • Avoids measuring the process instead of the result

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences