Skip to content
AI-grafen
EUniversityLab· about 60 min· server sandbox

Lab: an agent loop with tools, a budget and a trace

Build a ReAct loop (think → act → observe) with two tools, a budget cap and a trace you can read afterwards — against a deterministic policy instead of an LLM.

Theory

At each step the agent picks a tool and arguments based on the task and the observations so far, until it answers or the budget runs out. Without a stopping condition it runs away.

Sub-tasks

  1. tools — calculator(expr) (only digits and + - * / parentheses, otherwise an error) and lookup(key) against a small fact base.
  2. policy — policy(task, trace) → ('calculator', expr) | ('lookup', key) | ('answer', text) — rule-based.
  3. run_agent — run_agent(task, budget) runs the loop, logs every step in trace, stops at answer or budget → {answer, steps, trace, stopped_by}.

Passes when: success >= 0.8

The starter code

runs in an isolated sandbox on the server
import re

FACTS = {"sveriges huvudstad": "Stockholm", "antal planeter": "8", "pi": "3.14159", "ljusets hastighet km/s": "299792"}


def calculator(expr):
    # TODO: tillåt bara [0-9 +-*/(). ]; annars raise ValueError; returnera str(resultat)
    ...


def lookup(key):
    # TODO: FACTS.get(key.lower().strip(), "okänt")
    ...


def policy(task, trace):
    """Regelbaserad 'LLM': bestämmer nästa handling utifrån uppgiften och tidigare observationer."""
    # TODO:
    #  - om senaste observationen finns och uppgiften är ett rent räkneuttryck → ('answer', obs)
    #  - om uppgiften innehåller 'räkna' eller ett uttryck med siffror och operator → ('calculator', uttryck)
    #  - om uppgiften börjar med 'vad är' och matchar en nyckel i FACTS → ('lookup', nyckel), sedan ('answer', obs)
    #  - annars ('answer', 'vet inte')
    ...


def run_agent(task, budget=4):
    trace = []
    # TODO: loop max budget varv: action = policy(task, trace); utför; logga {"action", "arg", "observation"}; answer → return
    ...

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

≥ 80 % of the tasks are solved within 4 steps; an unsolvable task is stopped by the budget with stopped_by='budget'.

Common mistakes

  • eval() on a free-form expression — validate the characters first.
  • The loop has no budget check.
  • The trace does not log the observation.