Benchmarks · Lab: an agent loop with tools, a budget and a trace
Lab: an agent loop with tools, a budget and a trace — task-success
The measure: success (>= = higher is better). One entry per user (the best run).
| # | Name | Value | When |
|---|---|---|---|
| No entries yet. | |||
Submit a run
Sign in and run the lab in the sandbox.