EUniversityLab· about 60 min· server sandbox
Lab: an agent loop with tools, a budget and a trace
Build a ReAct loop (think → act → observe) with two tools, a budget cap and a trace you can read afterwards — against a deterministic policy instead of an LLM.
Teaches: The ReAct loop: think, act, observe
Requires: Tool usePython — classes and objects
Theory
At each step the agent picks a tool and arguments based on the task and the observations so far, until it answers or the budget runs out. Without a stopping condition it runs away.
Sub-tasks
- tools —
calculator(expr)(only digits and + - * / parentheses, otherwise an error) andlookup(key)against a small fact base. - policy —
policy(task, trace)→ ('calculator', expr) | ('lookup', key) | ('answer', text) — rule-based. - run_agent —
run_agent(task, budget)runs the loop, logs every step in trace, stops at answer or budget → {answer, steps, trace, stopped_by}.
Passes when: success >= 0.8
The starter code
runs in an isolated sandbox on the serverimport re
FACTS = {"sveriges huvudstad": "Stockholm", "antal planeter": "8", "pi": "3.14159", "ljusets hastighet km/s": "299792"}
def calculator(expr):
# TODO: tillåt bara [0-9 +-*/(). ]; annars raise ValueError; returnera str(resultat)
...
def lookup(key):
# TODO: FACTS.get(key.lower().strip(), "okänt")
...
def policy(task, trace):
"""Regelbaserad 'LLM': bestämmer nästa handling utifrån uppgiften och tidigare observationer."""
# TODO:
# - om senaste observationen finns och uppgiften är ett rent räkneuttryck → ('answer', obs)
# - om uppgiften innehåller 'räkna' eller ett uttryck med siffror och operator → ('calculator', uttryck)
# - om uppgiften börjar med 'vad är' och matchar en nyckel i FACTS → ('lookup', nyckel), sedan ('answer', obs)
# - annars ('answer', 'vet inte')
...
def run_agent(task, budget=4):
trace = []
# TODO: loop max budget varv: action = policy(task, trace); utför; logga {"action", "arg", "observation"}; answer → return
...
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
≥ 80 % of the tasks are solved within 4 steps; an unsolvable task is stopped by the budget with stopped_by='budget'.
Common mistakes
eval()on a free-form expression — validate the characters first.- The loop has no budget check.
- The trace does not log the observation.