Skip to content
AI-grafen
GFrontier LabEvals and benchmarks· about 180 min· fast-moving, sources checked often· verified 2026-09-20· EN

Build your own benchmark

Be able to design, validate and publish a benchmark with a clear measurement domain.

Prerequisites

Intuition

A benchmark is a measuring instrument. Like every instrument it has to have a defined measurement domain, a calibration and known sources of error.

Building it:

  1. The measurement domain (one sentence): «the ability to answer questions about Swedish upper-secondary mathematics with the correct final answer». Everything outside it is outside.
  2. The case construction: the source (written by experts / derived from licensed teaching material / synthetic + reviewed), the template, the difficulty levels, ≥ 300 cases. A key with annotator agreement (κ) and adjudication.
  3. The measure + the script: a deterministic parse, normalisation, one command that gives the figure.
  4. The validation: run chance, the majority, a simple model, two LLMs. A good benchmark: a spread between models, no ceiling, no cases all of them manage or none of them manage (unless deliberately).
  5. Contamination protection: a private test part, canary strings, the publication date, a licence forbidding training.
  6. Publication: a version number, a data card (the source, the licence, the bias, the intended use), the eval script, a leaderboard with the configuration per row.

AI-grafen exposes its own benchmarks (/benchmarks) in exactly this way: the labs' evals are the cases, the runs in the sandbox are the measurement.

Research

Known faults in published benchmarks: incorrect keys (MMLU has several per cent wrong answers in some subjects), ambiguous questions, contamination, cultural and linguistic bias, saturation (HumanEval), and the top of the leaderboard being driven by prompt tricks rather than model ability. The countermeasures in the literature: dynamic benchmarks (Dynabench), «living benchmarks» with rotation (LiveBench), private test sets with an eval server (SWE-bench Verified), pairwise arenas (Chatbot Arena) — each with weaknesses of its own (Elo drift, user selection).

A Swedish benchmark for AI learning would need: cases per level A–G, the measurement domain «explain / correct / tutor without giving the solution away», human calibration with teachers and pupils, and rotation every term.

Mastery means

  • Defines the measurement domain, the case construction and the key with annotator agreement
  • Validates the benchmark: the spread of difficulty, contamination protection, baselines
  • Publishes it with a licence, versioning and an eval script

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences