FAI engineeringEvals and benchmarks· about 90 min· fast-moving, sources checked often· verified 2026-09-20· EN
An LLM as a judge
Be able to use a model for assessment, measure the agreement with people and recognise the bias.
Prerequisites
- FEvals for language models and agentsrequired
Intuition
When the answers are free text, human assessment does not scale. An LLM as a judge assesses instead — but it has systematic errors that have to be handled:
| The bias | What it does | The remedy |
|---|---|---|
| Position | prefers the first (or the last) alternative | randomise the order, run both and average |
| Length | longer answers are judged better | explicit length neutrality in the rubric; check the correlation |
| Self-preference | prefers text from the same model family | use a different model as judge from the one that answered |
| Style | lists and headings are rewarded | a rubric that separates form from content |
| Leniency | nearly everything gets 4 out of 5 | pairwise instead of an absolute scale |
Pairwise comparison with a randomised order solves most of this at once.
Code
import random, numpy as np
RUBRIC = """Compare two answers to the same question. Judge the content ONLY: correctness first,
then completeness. Do NOT let yourself be affected by length, formatting or a confident tone.
Answer with exactly one of: A, B, EQUAL."""
def pairwise(llm, question, answer_a, answer_b, rng):
"""Run both orders — if the judge is inconsistent it counts as EQUAL."""
votes = []
for flip in (False, True):
x, y = (answer_b, answer_a) if flip else (answer_a, answer_b)
out = llm(f"{RUBRIC}\n\nQUESTION: {question}\n\nANSWER 1:\n{x}\n\nANSWER 2:\n{y}",
temperature=0).strip().upper()[:5]
if out.startswith("EQUAL"):
votes.append("EQUAL")
else:
first = out.startswith("A") or out.startswith("1")
votes.append(("B" if flip else "A") if first else ("A" if flip else "B"))
return votes[0] if votes[0] == votes[1] else "EQUAL"
def calibrate(judge_votes, human_votes):
"""The agreement between the judge and a person — always report this."""
pairs = [(j, h) for j, h in zip(judge_votes, human_votes) if h != "EQUAL"]
return sum(j == h for j, h in pairs) / max(len(pairs), 1)
# Below about 0.75 agreement: the rubric is not good enough. Rewrite it and measure again.
Calibration is not optional. Take 100 cases, have two people judge them, measure the judge's agreement — and report that figure alongside every result the judge produces. Without it nobody knows what the measure is worth.
Mastery means
- Builds a judge with a rubric and pairwise comparison
- Measures the agreement with human judgements
- Recognises and counteracts the judge's bias
Sign in to do the exercises and build your mastery up.
Sources
- arXiv — Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — arXiv (open access; licence per article)
- arXiv — Large Language Models are not Fair Evaluators — arXiv (open access; licence per article)