FAI engineeringAgents and tool use· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Evaluating agents
Be able to build task-based evals for agents with success criteria and cost.
Prerequisites
- EThe ReAct loop: think, act, observerequired
- FEvals for language models and agentsrequired
Intuition
Agents cannot be evaluated like chat answers. The question is not «did the answer sound good?» but «was the task solved?».
Build evals with a verifiable end state:
| The task | The verification |
|---|---|
| «book the meeting» | is the entry in the calendar with the right time? |
| «fix the bug» | does the test suite pass? |
| «find the price» | does the answer match the key? |
| «clean the data file» | does the file satisfy the schema? |
Run in a sandbox with a controlled starting state, let the agent work, inspect the end state with code. No judge is needed — it is binary.
Always report three numbers together: the solution rate, the cost per successful task, and the share aborted.
Code
from dataclasses import dataclass
from typing import Callable
@dataclass
class AgentCase:
id: str
task: str
start_state: Callable[[], object] # builds the sandbox
verify: Callable[[object], bool] # inspects the end state
def run_eval_suite(agent, cases: list[AgentCase], attempts: int = 3):
rows = []
for c in cases:
for k in range(attempts): # several attempts: agents are stochastic
world = c.start_state()
res = agent.run(c.task, world)
rows.append({"case": c.id, "attempt": k, "succeeded": bool(c.verify(world)),
"steps": res["steps"], "tokens": res["tokens"], "sec": res["sec"],
"status": res["status"]})
n = len(rows); ok = [r for r in rows if r["succeeded"]]
return {
"solution_rate": len(ok) / n,
"pass_at_1": sum(r["succeeded"] for r in rows if r["attempt"] == 0) / len(cases),
"cost_per_success": sum(r["tokens"] for r in rows) / max(len(ok), 1),
"mean_steps_successful": sum(r["steps"] for r in ok) / max(len(ok), 1),
"aborted": sum(r["status"] != "done" for r in rows) / n,
}
Two traps:
- Measuring the process. «Did the agent use the right tools in the right order?» sounds reasonable but punishes smarter solutions. Measure the end state.
- One attempt per case. Agents are stochastic;
pass@1andpass@3say different things and both are relevant.
Mastery means
- Builds task-based evals with a verifiable end state
- Measures the cost and the steps, not just success
- Avoids measuring the process instead of the result
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — arXiv (open access; licence per article)
- arXiv — AgentBench: Evaluating LLMs as Agents — arXiv (open access; licence per article)