Skip to content
AI-grafen
GFrontier LabLab· about 90 min· server sandbox

Lab: build and validate a benchmark

Build the tools that turn a set of cases into a benchmark — inter-annotator agreement, a contamination check against a corpus, validation against baselines (floor, ceiling, discrimination) and a data card — and run them on a given set of cases.

Theory

A benchmark is a measuring instrument: a measurement domain, an answer key with known reliability (κ), protection against contamination (n-gram overlap, canaries), and validation showing that it separates systems (no saturation, no floor). A data card documents the source, licence, construction and known biases.

Sub-tasks

  1. kappa — cohens_kappa(a, b) for two label lists (strings) of equal length. load_jsonl(path).
  2. contamination — ngrams(text, n) → set of n-grams (words, lower case, only \w+); contaminated(cases, corpus_text, n=8) → list of case ids whose question has at least one n-gram in the corpus.
  3. validation — validate(cases, results) where results is {system: {case_id: bool}} → {"floor": [ids nobody passes], "ceiling": [ids everybody passes], "discrimination": fraction of cases where the systems differ, "accuracy": {system: fraction}}.
  4. data card — datacard(cases, meta) → markdown string with the headings "# <title>", "## Mätdomän", "## Källa och licens", "## Konstruktion", "## Kända skevheter", "## Statistik" (the tests check these Swedish headings: measurement domain, source and licence, construction, known biases, statistics — number of cases, distribution by difficulty).

Passes when: ok >= 1

The starter code

runs in an isolated sandbox on the server
"""Benchmark-verktyg. Fyll i funktionerna."""
import json
import re


def load_jsonl(path):
    return [json.loads(l) for l in open(path, encoding="utf-8") if l.strip()]


def cohens_kappa(a, b):
    """Cohens kappa för två lika långa listor med etiketter."""
    # TODO
    raise NotImplementedError


def ngrams(text, n=8):
    """Mängd av n-gram över ord (gemener, bara \\w+)."""
    # TODO
    raise NotImplementedError


def contaminated(cases, corpus_text, n=8):
    """Lista av case-id (i indataordning) vars 'question' delar minst ett n-gram med korpusen."""
    # TODO
    raise NotImplementedError


def validate(cases, results):
    """results: {system: {case_id: bool}}. Returnerar {"floor": [...], "ceiling": [...], "discrimination": float, "accuracy": {system: float}}."""
    # TODO
    raise NotImplementedError


def datacard(cases, meta):
    """Markdown med rubrikerna: '# <title>', '## Mätdomän', '## Källa och licens', '## Konstruktion', '## Kända skevheter', '## Statistik'."""
    # TODO
    raise NotImplementedError

You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.

Try the diagnosticCreate a free account

Expected results

κ between the two annotators ≈ 0.7; exactly two cases are contaminated by the corpus (c07, c14); the validation finds one ceiling case and one floor case; the data card contains all six headings. The eval sets ok = 1 when everything is right.

Common mistakes

  • κ: pe should sum the product of the marginals per label, not just the majority class.
  • n-grams: normalise to lower case and remove punctuation before splitting — otherwise you miss overlap.
  • discrimination: a case discriminates when at least one system is right and at least one is wrong.
  • The data card: the headings must be exactly the ones given (the tests look for them).