FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
Human evaluation and annotator agreement
Be able to design human evaluation with rubrics and measure the agreement.
Prerequisites
- EAnnotation and labelling of datarequired
- EHypothesis testing and p-valuesrequired
Intuition
Automatic metrics are proxies. In the end people have to judge whether the answers are good — and that judgement has to be designed as carefully as an experiment.
- What is judged: an absolute score (1–5 per criterion) or pairwise (A vs B) — pairwise is more reliable and simpler.
- The rubric: definitions per score with examples; a pilot; revise.
- The sample: random from the production distribution, stratified by difficulty; n ≥ 200 to see 5 pp.
- Blinding: the judge does not know which model; the order is randomised.
- The judges: ≥ 2 per case, measure κ/α; domain experts for domain questions.
- The cost: 200 cases × 2 judges × 3 min ≈ 20 hours. Plan for it.
The result is used twice: as the answer key now, and to calibrate an LLM judge so that the next 10 000 cases can be assessed automatically with a known margin of error.
Code
import random, csv
def build_rating_task(cases, answers_a, answers_b, seed=0):
"""Creates blinded, randomised pairwise tasks. The key to who was A/B is saved separately."""
rng = random.Random(seed); tasks, key = [], []
for i, (c, a, b) in enumerate(zip(cases, answers_a, answers_b)):
flip = rng.random() < 0.5
tasks.append({"id": i, "question": c, "answer_1": b if flip else a, "answer_2": a if flip else b})
key.append({"id": i, "answer_1_is": "B" if flip else "A"})
return tasks, key
# after the rating: unrandomise and count
def results(ratings, key):
k = {n["id"]: n["answer_1_is"] for n in key}; wins = {"A": 0, "B": 0, "equal": 0}
for r in ratings: # r = {"id", "choice": "1"|"2"|"equal"}
if r["choice"] == "equal": wins["equal"] += 1
else: wins[k[r["id"]] if r["choice"] == "1" else ("B" if k[r["id"]] == "A" else "A")] += 1
return wins
Report: the win share with a CI, κ between the judges, and the share of «equal».
Mastery means
- Designs human evaluation with a rubric, a sample and blinding
- Measures annotator agreement and reports it
- Uses human data to calibrate automatic judges
Sign in to do the exercises and build your mastery up.
Sources
- Wikipedia — Inter-rater reliability (CC BY-SA 4.0) — CC BY-SA 4.0
- arXiv — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arXiv (open access; licence per article)