Skip to content
AI-grafen
EUniversityAgents and tool use· about 60 min· evolving, reviewed regularly· verified 2026-09-20· EN

The ReAct loop: think, act, observe

Be able to implement a ReAct agent and read its trace.

Prerequisites

Intuition

ReAct = Reasoning + Acting. The model alternates between thinking and acting, in a trace that can be read:

Thought:     I need to know how many cases are open.
Action:      search_cases(status="open")
Observation: 47 cases
Thought:     Now I need how many of them are older than a week.
Action:      search_cases(status="open", older_than_days=7)
Observation: 12 cases
Thought:     I have what I need.
Answer:      12 of the 47 open cases are older than a week.

The point of the explicit thoughts is not magic — it is readability and debugging. When the agent goes wrong you see exactly at which step the reasoning came off the rails.

Code

SYSTEM = """You solve tasks step by step. Answer in exactly this format:
Thought: <your reasoning>
Action: <tool>(<arguments>)
… or when you are finished:
Thought: <reasoning>
Answer: <the final answer>
Use only these tools: {tools}"""

import re

def react(goal, llm, tools, max_steps=8):
    trace, msgs = [], [{"role": "system", "content": SYSTEM.format(tools=tools.descriptions())},
                       {"role": "user", "content": goal}]
    for step in range(max_steps):
        out = llm(msgs, stop=["Observation:"], temperature=0)
        trace.append(out)
        if "Answer:" in out:
            return {"status": "done", "answer": out.split("Answer:", 1)[1].strip(), "trace": trace}
        m = re.search(r"Action:\s*(\w+)\((.*?)\)\s*$", out, re.S | re.M)
        if not m:
            obs = "Error: could not parse the action. Use the format Action: tool(arguments)."
        else:
            obs = tools.run(m.group(1), m.group(2))
        trace.append(f"Observation: {obs}")
        msgs.append({"role": "assistant", "content": out})
        msgs.append({"role": "user", "content": f"Observation: {obs}"})
    return {"status": "budget_exhausted", "answer": None, "trace": trace}

Three things that make the loop work in practice: the stop sequence so that the model does not invent its own observations, error messages returned as observations (the model can correct itself), and a hard cap on the number of steps.

Mastery means

  • Implements a ReAct loop
  • Reads and debugs an agent trace
  • Sets stopping conditions

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences