The ReAct loop: think, act, observe
Be able to implement a ReAct agent and read its trace.
Prerequisites
- EAgents — plan, act, observerequired
Intuition
ReAct = Reasoning + Acting. The model alternates between thinking and acting, in a trace that can be read:
Thought: I need to know how many cases are open.
Action: search_cases(status="open")
Observation: 47 cases
Thought: Now I need how many of them are older than a week.
Action: search_cases(status="open", older_than_days=7)
Observation: 12 cases
Thought: I have what I need.
Answer: 12 of the 47 open cases are older than a week.
The point of the explicit thoughts is not magic — it is readability and debugging. When the agent goes wrong you see exactly at which step the reasoning came off the rails.
Code
SYSTEM = """You solve tasks step by step. Answer in exactly this format:
Thought: <your reasoning>
Action: <tool>(<arguments>)
… or when you are finished:
Thought: <reasoning>
Answer: <the final answer>
Use only these tools: {tools}"""
import re
def react(goal, llm, tools, max_steps=8):
trace, msgs = [], [{"role": "system", "content": SYSTEM.format(tools=tools.descriptions())},
{"role": "user", "content": goal}]
for step in range(max_steps):
out = llm(msgs, stop=["Observation:"], temperature=0)
trace.append(out)
if "Answer:" in out:
return {"status": "done", "answer": out.split("Answer:", 1)[1].strip(), "trace": trace}
m = re.search(r"Action:\s*(\w+)\((.*?)\)\s*$", out, re.S | re.M)
if not m:
obs = "Error: could not parse the action. Use the format Action: tool(arguments)."
else:
obs = tools.run(m.group(1), m.group(2))
trace.append(f"Observation: {obs}")
msgs.append({"role": "assistant", "content": out})
msgs.append({"role": "user", "content": f"Observation: {obs}"})
return {"status": "budget_exhausted", "answer": None, "trace": trace}
Three things that make the loop work in practice: the stop sequence so that the model does not invent its own observations, error messages returned as observations (the model can correct itself), and a hard cap on the number of steps.
Mastery means
- Implements a ReAct loop
- Reads and debugs an agent trace
- Sets stopping conditions
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — ReAct: Synergizing Reasoning and Acting in Language Models — arXiv (open access; licence per article)