Skip to content
AI-grafen
FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN

An LLM as a judge

Be able to use a model for assessment, measure the agreement with people and recognise the bias.

Prerequisites

Intuition

When the answers are free text, human assessment does not scale. An LLM as a judge assesses instead — but it has systematic errors that have to be handled:

The biasWhat it doesThe remedy
Positionprefers the first (or the last) alternativerandomise the order, run both and average
Lengthlonger answers are judged betterexplicit length neutrality in the rubric; check the correlation
Self-preferenceprefers text from the same model familyuse a different model as judge from the one that answered
Stylelists and headings are rewardeda rubric that separates form from content
Leniencynearly everything gets 4 out of 5pairwise instead of an absolute scale

Pairwise comparison with a randomised order solves most of this at once.

Code

import random, numpy as np

RUBRIC = """Compare two answers to the same question. Judge the content ONLY: correctness first,
then completeness. Do NOT let yourself be affected by length, formatting or a confident tone.
Answer with exactly one of: A, B, EQUAL."""

def pairwise(llm, question, answer_a, answer_b, rng):
    """Run both orders — if the judge is inconsistent it counts as EQUAL."""
    votes = []
    for flip in (False, True):
        x, y = (answer_b, answer_a) if flip else (answer_a, answer_b)
        out = llm(f"{RUBRIC}\n\nQUESTION: {question}\n\nANSWER 1:\n{x}\n\nANSWER 2:\n{y}",
                  temperature=0).strip().upper()[:5]
        if out.startswith("EQUAL"):
            votes.append("EQUAL")
        else:
            first = out.startswith("A") or out.startswith("1")
            votes.append(("B" if flip else "A") if first else ("A" if flip else "B"))
    return votes[0] if votes[0] == votes[1] else "EQUAL"

def calibrate(judge_votes, human_votes):
    """The agreement between the judge and a person — always report this."""
    pairs = [(j, h) for j, h in zip(judge_votes, human_votes) if h != "EQUAL"]
    return sum(j == h for j, h in pairs) / max(len(pairs), 1)

# Below about 0.75 agreement: the rubric is not good enough. Rewrite it and measure again.

Calibration is not optional. Take 100 cases, have two people judge them, measure the judge's agreement — and report that figure alongside every result the judge produces. Without it nobody knows what the measure is worth.

Mastery means

  • Builds a judge with a rubric and pairwise comparison
  • Measures the agreement with human judgements
  • Recognises and counteracts the judge's bias

Sign in to do the exercises and build your mastery up.

Sources

All the sources and licences