Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

Human evaluation and annotator agreement

Be able to design human evaluation with rubrics and measure the agreement.

Prerequisites

Intuition

Automatic metrics are proxies. In the end people have to judge whether the answers are good — and that judgement has to be designed as carefully as an experiment.

  • What is judged: an absolute score (1–5 per criterion) or pairwise (A vs B) — pairwise is more reliable and simpler.
  • The rubric: definitions per score with examples; a pilot; revise.
  • The sample: random from the production distribution, stratified by difficulty; n ≥ 200 to see 5 pp.
  • Blinding: the judge does not know which model; the order is randomised.
  • The judges: ≥ 2 per case, measure κ/α; domain experts for domain questions.
  • The cost: 200 cases × 2 judges × 3 min ≈ 20 hours. Plan for it.

The result is used twice: as the answer key now, and to calibrate an LLM judge so that the next 10 000 cases can be assessed automatically with a known margin of error.

Code

import random, csv

def build_rating_task(cases, answers_a, answers_b, seed=0):
    """Creates blinded, randomised pairwise tasks. The key to who was A/B is saved separately."""
    rng = random.Random(seed); tasks, key = [], []
    for i, (c, a, b) in enumerate(zip(cases, answers_a, answers_b)):
        flip = rng.random() < 0.5
        tasks.append({"id": i, "question": c, "answer_1": b if flip else a, "answer_2": a if flip else b})
        key.append({"id": i, "answer_1_is": "B" if flip else "A"})
    return tasks, key

# after the rating: unrandomise and count
def results(ratings, key):
    k = {n["id"]: n["answer_1_is"] for n in key}; wins = {"A": 0, "B": 0, "equal": 0}
    for r in ratings:                          # r = {"id", "choice": "1"|"2"|"equal"}
        if r["choice"] == "equal": wins["equal"] += 1
        else: wins[k[r["id"]] if r["choice"] == "1" else ("B" if k[r["id"]] == "A" else "A")] += 1
    return wins

Report: the win share with a CI, κ between the judges, and the share of «equal».

Mastery means

  • Designs human evaluation with a rubric, a sample and blinding
  • Measures annotator agreement and reports it
  • Uses human data to calibrate automatic judges

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences