Skip to content
AI-grafen

Benchmarks · Lab: an agent loop with tools, a budget and a trace

Lab: an agent loop with tools, a budget and a trace — task-success

The measure: success (>= = higher is better). One entry per user (the best run).

#NameValueWhen
No entries yet.

Submit a run

Sign in and run the lab in the sandbox.