Multi-agent systems
Be able to build collaborating agents with roles and measure whether it beats a single agent.
Prerequisites
- FAgent architecturesrequired
Intuition
Multi-agent means several LLM instances with different roles and prompts passing work between them: a planner, an executor, a reviewer, a compiler.
The argument for: every role gets a focused prompt and its own context, the reviewer sees the work with «fresh eyes», and the parts can be run in parallel.
The argument against: every handover happens in natural language and loses precision; the cost is multiplied; the debugging becomes harder; and the advantages often come from the task having been split up — not from several models «collaborating».
The rule: always build the single-agent variant first and measure it. Multi-agent has to prove that it is enough better to justify 3–10× the cost.
Code
from dataclasses import dataclass
@dataclass
class Role:
name: str
system: str
tools: list[str]
ROLES = [
Role("planner", "Break the task down into 3-5 subgoals. Carry out nothing yourself.", []),
Role("executor", "Carry out ONE subgoal with the tools. Report the result and the uncertainty.", ["search", "read"]),
Role("reviewer", "Review the result against the subgoal. Answer APPROVED or list the shortcomings.", []),
]
def multi_agent(task, llm, tools, max_rounds=3):
plan = llm.run(ROLES[0], task)
results, log = [], []
for subgoal in plan["subgoals"]:
for rnd in range(max_rounds):
r = llm.run(ROLES[1], subgoal, tools=tools, previous=results)
g = llm.run(ROLES[2], {"subgoal": subgoal, "result": r})
log.append({"subgoal": subgoal["id"], "round": rnd, "review": g["status"]})
if g["status"] == "APPROVED":
results.append(r); break
subgoal["feedback"] = g["shortcomings"] # the executor gets concrete criticism
else:
results.append({**r, "uncertain": True}) # gave up after max_rounds
return {"results": results, "log": log}
The measurement that decides whether it was worth it — run both on the same eval suite:
| the solution rate | tokens/task | the time | |
|---|---|---|---|
| one agent (ReAct) | 0.71 | 18 k | 22 s |
| multi-agent | 0.78 | 96 k | 71 s |
+7 percentage points for 5.3× the cost. Whether that is worth it depends on the value of the task — but the decision should be made on figures, not on the architecture sounding sophisticated.
Mastery means
- Builds collaborating agents with roles
- Measures whether the division beats a single agent
- Recognises when the cost is not justified
Sign in to do the exercises and build your mastery up.
Sources
- Anthropic — Building effective agents — free to read
- arXiv — AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation — arXiv (open access; licence per article)