Lab: build and validate a benchmark
Build the tools that turn a set of cases into a benchmark — inter-annotator agreement, a contamination check against a corpus, validation against baselines (floor, ceiling, discrimination) and a data card — and run them on a given set of cases.
Theory
A benchmark is a measuring instrument: a measurement domain, an answer key with known reliability (κ), protection against contamination (n-gram overlap, canaries), and validation showing that it separates systems (no saturation, no floor). A data card documents the source, licence, construction and known biases.
Sub-tasks
- kappa —
cohens_kappa(a, b)for two label lists (strings) of equal length.load_jsonl(path). - contamination —
ngrams(text, n)→ set of n-grams (words, lower case, only \w+);contaminated(cases, corpus_text, n=8)→ list of case ids whose question has at least one n-gram in the corpus. - validation —
validate(cases, results)where results is {system: {case_id: bool}} → {"floor": [ids nobody passes], "ceiling": [ids everybody passes], "discrimination": fraction of cases where the systems differ, "accuracy": {system: fraction}}. - data card —
datacard(cases, meta)→ markdown string with the headings "# <title>", "## Mätdomän", "## Källa och licens", "## Konstruktion", "## Kända skevheter", "## Statistik" (the tests check these Swedish headings: measurement domain, source and licence, construction, known biases, statistics — number of cases, distribution by difficulty).
Passes when: ok >= 1
The starter code
runs in an isolated sandbox on the server"""Benchmark-verktyg. Fyll i funktionerna."""
import json
import re
def load_jsonl(path):
return [json.loads(l) for l in open(path, encoding="utf-8") if l.strip()]
def cohens_kappa(a, b):
"""Cohens kappa för två lika långa listor med etiketter."""
# TODO
raise NotImplementedError
def ngrams(text, n=8):
"""Mängd av n-gram över ord (gemener, bara \\w+)."""
# TODO
raise NotImplementedError
def contaminated(cases, corpus_text, n=8):
"""Lista av case-id (i indataordning) vars 'question' delar minst ett n-gram med korpusen."""
# TODO
raise NotImplementedError
def validate(cases, results):
"""results: {system: {case_id: bool}}. Returnerar {"floor": [...], "ceiling": [...], "discrimination": float, "accuracy": {system: float}}."""
# TODO
raise NotImplementedError
def datacard(cases, meta):
"""Markdown med rubrikerna: '# <title>', '## Mätdomän', '## Källa och licens', '## Konstruktion', '## Kända skevheter', '## Statistik'."""
# TODO
raise NotImplementedError
You write the code; tests you cannot see decide whether it holds up. Create a free account to run the lab.
Try the diagnosticCreate a free accountExpected results
κ between the two annotators ≈ 0.7; exactly two cases are contaminated by the corpus (c07, c14); the validation finds one ceiling case and one floor case; the data card contains all six headings. The eval sets ok = 1 when everything is right.
Common mistakes
- κ: pe should sum the product of the marginals per label, not just the majority class.
- n-grams: normalise to lower case and remove punctuation before splitting — otherwise you miss overlap.
- discrimination: a case discriminates when at least one system is right and at least one is wrong.
- The data card: the headings must be exactly the ones given (the tests look for them).