Budget, stopping conditions and cost control
Be able to set budgets and stopping conditions so that an agent never runs away.
Prerequisites
- EAgents — plan, act, observerequired
- ECost modelling for LLM systemsrequired
Intuition
An agent that loops costs money every round. Unlike a single call there is no natural limit — the agent decides for itself when it is finished, and sometimes it never does.
Five independent stopping conditions. All of them are needed, because they catch different kinds of breakdown:
| The condition | Catches |
|---|---|
| Max steps | loops |
| Max time | slow tools, hangs |
| Max cost | expensive model calls |
| Max tool calls per type | an agent calling the same thing over and over |
| No progress | the agent repeating itself without getting anywhere |
The last is the hardest to implement and the most valuable: an agent can be within every numerical limit and still not be doing anything meaningful.
Formal
Detecting a lack of progress. Three signals, in order of how easy they are to implement:
| The signal | How |
|---|---|
| An identical tool call is repeated | hash the arguments and count |
| The observations stop changing | hash the state after every step |
| No new information | compare the entropy or the novelty of the observations |
The first goes a long way: if the same call with the same arguments is made three times the agent is stuck.
A budget that degrades gracefully. An agent that simply stops at 100 % of the budget gives the user nothing. Better to step down:
| The budget used | The behaviour |
|---|---|
| < 50 % | normal |
| 50–75 % | switch to a cheaper model for routine steps |
| 75–90 % | «finish what you are doing, no new subgoals» |
| 90–100 % | summarise what has been done and what remains |
| 100 % | stop with a usable partial delivery |
The fourth row is the one that makes the difference: an agent that says «I managed A and B, C remains and here is what I know» is useful. One that is merely cut off is not.
Cost attribution per step makes improvement possible:
| What is logged | Why |
|---|---|
| Tokens in and out per call | where the cost is |
| Which tool and for how long | slow tools |
| Which subgoal the step belonged to | which part of the task is expensive |
| Whether the step gave new information | inefficiency |
The most common cost driver in agents is not the model but the context. The history grows with every step, and since the whole history is sent along every time the cost grows quadratically with the number of steps. An agent with 30 steps can cost more than ten times one with 10 steps, not three times.
The countermeasures: summarise old steps, throw away tool output that is no longer needed, and use prompt caching on the fixed part.
Set the budget per task, not per call. The question «what is this task worth?» has an answer; «what is a call worth?» does not.
Code
import hashlib, json, time
from collections import Counter
from dataclasses import dataclass, field
@dataclass
class Budget:
max_steps: int = 30
max_seconds: float = 300.0
max_kr: float = 5.0
max_per_tool: int = 8
max_repeats: int = 3
steps: int = 0
kr: float = 0.0
start: float = field(default_factory=time.perf_counter)
tool_counts: Counter = field(default_factory=Counter)
call_hashes: Counter = field(default_factory=Counter)
state_hashes: list = field(default_factory=list)
log: list = field(default_factory=list)
def used(self) -> float:
return max(self.steps / self.max_steps,
(time.perf_counter() - self.start) / self.max_seconds,
self.kr / self.max_kr)
def mode(self) -> str:
u = self.used()
if u < 0.50: return "normal"
if u < 0.75: return "save" # a cheaper model for routine steps
if u < 0.90: return "finish" # no new subgoals
if u < 1.00: return "summarise"
return "stop"
def check(self, tool: str, arguments: dict, state: str):
h = hashlib.sha256(
f"{tool}:{json.dumps(arguments, sort_keys=True)}".encode()).hexdigest()[:16]
self.call_hashes[h] += 1
if self.call_hashes[h] > self.max_repeats:
return False, f"the same call to {tool} repeated {self.call_hashes[h]} times"
if self.tool_counts[tool] >= self.max_per_tool:
return False, f"the maximum number of calls to {tool}"
self.state_hashes.append(state)
if len(self.state_hashes) >= 4 and len(set(self.state_hashes[-4:])) == 1:
return False, "the state has not changed for four steps"
if self.mode() == "stop":
return False, "the budget has run out"
return True, self.mode()
def register(self, tool, in_tok, out_tok, price_in, price_out, new_information: bool):
cost = (in_tok * price_in + out_tok * price_out) / 1e6
self.steps += 1
self.kr += cost
self.tool_counts[tool] += 1
self.log.append({"step": self.steps, "tool": tool, "kr": round(cost, 4),
"in_tok": in_tok, "new_information": new_information})
def report(self):
wasted = sum(1 for r in self.log if not r["new_information"])
return {"steps": self.steps, "kr": round(self.kr, 3),
"seconds": round(time.perf_counter() - self.start, 1),
"most_expensive_tools": self.tool_counts.most_common(3),
"share_of_steps_without_new_information": round(wasted / max(self.steps, 1), 3),
"context_growth": [r["in_tok"] for r in self.log]}
# The context grows quadratically if the history is sent along every time
def cost_without_compression(steps, tokens_per_step=800, price_in=30.0):
return sum((i * tokens_per_step) * price_in / 1e6 for i in range(1, steps + 1))
for n in (10, 20, 30):
print(f"{n:>2} steps: {cost_without_compression(n):.3f} kr")
# 10 steps: 1.320 kr
# 20 steps: 5.040 kr
# 30 steps: 11.160 kr ← three times as many steps, eight times as expensive
The last line printed is the most important number in the node: the cost grows with the square of the number of steps if the history is not compressed.
Mastery means
- Sets several independent stopping conditions
- Detects loops and inefficiency
- Designs a budget that degrades gracefully
Sign in to do the exercises and build your mastery up.
Sources
- OWASP Top 10 for LLM Applications — CC BY-SA 4.0
- Anthropic — Tool use — documentation, free to read
- arXiv — WebArena: A Realistic Web Environment for Building Autonomous Agents — arXiv (open access; licence per article)