Skip to content
AI-grafen

Benchmarks · Lab: build an eval harness

Lab: build an eval harness — baseline-acc

The measure: accuracy (>= = higher is better). One entry per user (the best run).

#NameValueWhen
No entries yet.

Submit a run

Sign in and run the lab in the sandbox.