Skip to content
AI-grafen
FAI engineeringAgents and tool use· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Multi-agent systems

Be able to build collaborating agents with roles and measure whether it beats a single agent.

Prerequisites

Intuition

Multi-agent means several LLM instances with different roles and prompts passing work between them: a planner, an executor, a reviewer, a compiler.

The argument for: every role gets a focused prompt and its own context, the reviewer sees the work with «fresh eyes», and the parts can be run in parallel.

The argument against: every handover happens in natural language and loses precision; the cost is multiplied; the debugging becomes harder; and the advantages often come from the task having been split up — not from several models «collaborating».

The rule: always build the single-agent variant first and measure it. Multi-agent has to prove that it is enough better to justify 3–10× the cost.

Code

from dataclasses import dataclass

@dataclass
class Role:
    name: str
    system: str
    tools: list[str]

ROLES = [
    Role("planner", "Break the task down into 3-5 subgoals. Carry out nothing yourself.", []),
    Role("executor", "Carry out ONE subgoal with the tools. Report the result and the uncertainty.", ["search", "read"]),
    Role("reviewer", "Review the result against the subgoal. Answer APPROVED or list the shortcomings.", []),
]

def multi_agent(task, llm, tools, max_rounds=3):
    plan = llm.run(ROLES[0], task)
    results, log = [], []
    for subgoal in plan["subgoals"]:
        for rnd in range(max_rounds):
            r = llm.run(ROLES[1], subgoal, tools=tools, previous=results)
            g = llm.run(ROLES[2], {"subgoal": subgoal, "result": r})
            log.append({"subgoal": subgoal["id"], "round": rnd, "review": g["status"]})
            if g["status"] == "APPROVED":
                results.append(r); break
            subgoal["feedback"] = g["shortcomings"]   # the executor gets concrete criticism
        else:
            results.append({**r, "uncertain": True})  # gave up after max_rounds
    return {"results": results, "log": log}

The measurement that decides whether it was worth it — run both on the same eval suite:

the solution ratetokens/taskthe time
one agent (ReAct)0.7118 k22 s
multi-agent0.7896 k71 s

+7 percentage points for 5.3× the cost. Whether that is worth it depends on the value of the task — but the decision should be made on figures, not on the architecture sounding sophisticated.

Mastery means

  • Builds collaborating agents with roles
  • Measures whether the division beats a single agent
  • Recognises when the cost is not justified

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences