An annotated companion · AI Primer

Judging LLM-as-a-Judge, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is distributed under arXiv's standard licence, so its tables are not reproduced; key numbers appear in this page's own tables and are attributed. The diagrams and demos are original. Read the original alongside: every section links to it.

How to read this page

  • Any dotted word explains itself when you hover, tab to, or tap it, and so does every symbol in every equation.
  • Two demos: a position-bias simulator (judge twice with the answers swapped) and an agreement calculator (how much of an agreement rate is real).

Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. A calibrated judge in code is in the evaluation lesson.

Abstract

“Our results reveal that strong LLM judges like GPT-4 can match both controlled and crowdsourced human preferences well, achieving over 80% agreement, the same level of agreement between humans.”Zheng et al. (2023), Abstract. Read the original

Everyday picture

A cooking competition needs judges. Human judges are the gold standard but slow and expensive. Could an experienced chef grade thousands of dishes instead? Only if we first check how often the chef agrees with the human panel, and what quirks the chef has, such as always preferring the first plate tasted or the biggest portion. This paper does that check for LLM judges.

What the paper claims

  • A strong model (GPT-4) agrees with expert humans over 80% of the time on which of two chatbot answers is better: about as often as two humans agree with each other.
  • LLM judges have measurable biases: position (favouring the first or second answer), verbosity (favouring longer answers), possibly favouring their own answers, and trouble grading maths and reasoning.
  • Simple fixes help: swap the answer order and require a consistent verdict, show a reference answer, or ask the judge to solve the problem first.
  • Two public benchmarks: MT-Bench (80 two-turn questions) and Chatbot Arena (crowd-sourced head-to-head votes).

Why it matters today

Automated grading of open-ended output is now routine in building AI systems. This paper supplies the vocabulary (pairwise, single-answer, reference-guided; position and verbosity bias) and, more importantly, the habit: check the judge against humans before trusting it.

1 Introduction · original

Everyday picture

A multiple-choice exam can tell whether a student knows facts. It cannot tell whether they write a helpful, well-organised reply to a messy, open-ended request. Chatbots trained to be helpful were clearly preferred by people over their base models, yet standard benchmarks such as MMLU barely told them apart.

Tiny example

In the paper's own comparison, LLaMA-13B (a base model) and Vicuna-13B (the same model fine-tuned on conversations) score similarly on knowledge benchmarks, but on an open-ended follow-up request people strongly prefer Vicuna. The benchmark measured what the model knows; people care about how it helps.

2 MT-Bench and Chatbot Arena · original

MT-BenchChatbot Arena
What80 hand-written questions, each with a follow-up (2 turns)A website where people chat with two anonymous models at once and vote for the better reply
Coverage8 categories × 10: writing, roleplay, extraction, reasoning, maths, coding, STEM knowledge, humanities knowledgeWhatever real users ask
Judges in the paper58 expert labellers (mostly graduate students) and LLM judgesAbout 30K crowd votes in a month; a 3K sample was compared with LLM judges
StrengthControlled and repeatableReal use, at scale

An example of the two-turn style (written for this page): “Write a two-sentence product description for a travel mug” followed by “Now rewrite it for a ten-year-old, still in two sentences”. The second turn tests whether the model follows a new constraint while keeping the earlier content.

3 LLM as a judge · original

Hover or tap the parts of the diagram to compare the three set-ups.

pairwise single answer reference-guided question answer A answer B judge A / B / tie question answer judge score 1 to 10 question reference answer(s) judge verdict or score

Hover or tap a judge box to compare the set-ups.

The three LLM-as-a-judge set-ups described in §3.1 of Zheng et al. (2023), drawn for this page.

Reading it: each column is one set-up, read top to bottom. All three show the judge the question. Pairwise shows two answers and asks which is better. Single-answer shows one and asks for a score. Reference-guided adds a correct solution, which matters for maths and reasoning where the judge might not know the answer itself. The biases in §3.3 attach to particular columns: position bias only exists when there are two answers side by side.

3.1 Three kinds of judge · original

Everyday picture

A wine taster can compare two glasses side by side, score each glass alone, or taste against a known reference bottle. Side by side notices subtle differences but needs many comparisons; scoring alone is quick but drifts from day to day.

The cost of comparing everyone with everyone

In words: “every model must be compared with every other; each of m models meets m − 1 others, and each pairing is counted twice that way, so halve it.”

With the numbers: the paper's 6 models need 6 × 5 / 2 = 15 pairings per question; 20 models would need 190. Single-answer grading needs just one judgement per model per question.

In Python:

# models being compared
m = 6
# pairs = C(m, 2) = m(m − 1) / 2
m * (m - 1) // 2  # → 15
m = 20
m * (m - 1) // 2  # → 190

The trade-off, as the paper puts it: pairwise comparison scales poorly as models are added, while single-answer grading can miss subtle differences and its absolute scores shift more when the judge model changes.

3.3 Biases and limits · original

“Position bias is when an LLM exhibits a propensity to favor certain positions over others.”Zheng et al. (2023), §3.3

Position bias

Everyday picture: a talent show where the first act always seems to win. The test: for each first-turn MT-Bench question, two similar answers were generated, then each judge compared them twice, once in each order. A consistent judge picks the same answer both times.

Selected numbers from Zheng et al. (2023), Table 2 (default prompt). Summarised with attribution.
JudgeConsistent when swappedFavoured whichever came firstFavoured whichever came second
Claude-v123.8%75.0%0.0%
GPT-3.546.2%50.0%1.2%
GPT-465.0%30.0%5.0%

The answers here were deliberately very similar, so this is a hard test; the paper notes position bias is milder when answers differ more. Renaming the assistants showed that Claude-v1 was also partly biased towards the name “Assistant A”.

Verbosity bias

Everyday picture: an essay marker who rewards length. The test (a “repetitive list” attack): take 23 answers that contain a numbered list, have a model rephrase the list, and prepend the rephrased copy, so the answer is twice as long with no new information. The attack succeeds if the judge prefers the padded version.

Reading it: each bar is how often a judge fell for the padding (numbers from Zheng et al., 2023, Table 3). Two of the three judges preferred the padded, repetitive answer about nine times in ten. GPT-4 resisted most of the time, but not always. All three correctly called two identical answers a tie, so the problem is specifically being impressed by length.

Self-enhancement bias

Do judges favour their own answers? Compared with humans, GPT-4 gave itself about a 10% higher win rate and Claude-v1 gave itself about 25% more, but GPT-3.5 did not favour itself, and judges also favoured some other models. With limited data the paper does not conclude that the bias is real, only that it may be.

Grading maths and reasoning

A judge can be misled by the answers it is grading, even on problems it can solve on its own. On 10 maths questions (each judged in both orders, so 20 judgements), GPT-4 called an incorrect answer correct 14 times with the default prompt.

Try it: position bias, and the swap fix

A simulated judge compares 1,000 answer pairs twice, once in each order. Each pair has a true quality difference; the judge tends to prefer the better answer (sensitivity) but also leans towards whichever it sees first (bias). The model and its numbers are illustrative, not fitted to the paper's data.

In words: “the judge's chance of picking the first answer rises with how much better the first answer really is, plus a constant nudge towards whatever is shown first.”

With the numbers: with k = 4 and β = 1, two equally good answers (δ = 0) give σ(1) = 0.73: the first-shown answer wins 73% of the time for no reason. If the first answer is clearly worse (δ = −0.5), σ(−2 + 1) = σ(−1) = 0.27.

In Python:

import math
# σ squashes any number into (0, 1)
def sigma(z):
    return 1 / (1 + math.exp(-z))
# sensitivity k, lean towards first position β
k, beta = 4, 1
# two equally good answers
delta = 0
round(sigma(k * delta + beta), 2)  # → 0.73
# the first answer is clearly worse
delta = -0.5
round(sigma(k * delta + beta), 2)  # → 0.27

Reading it: the bars show three rates over the 1,000 simulated pairs. Consistent is how often the judge picks the same underlying answer in both orders, the paper's consistency measure. Single-order errors is how often a one-shot judgement picks the truly worse answer. Swap-rule errors applies the paper's conservative fix: declare a winner only if both orders agree, otherwise call it a tie, and count only wrong winners as errors. Push β up: consistency falls, single-order errors rise, and the swap rule converts most of them into ties. That is its whole purpose.

3.4 Fixes · original

FixHowEffect reported in the paper
Swap positionsJudge twice with the order swapped; win only if both agree, otherwise tieRemoves the effect of position bias from verdicts (used in all later experiments)
Few-shot judgeShow 3 example judgements (A better, B better, tie)GPT-4 consistency 65.0% → 77.5%, but the prompt costs about 4 times as much
Chain-of-thought judgeAsk the judge to solve the problem itself before gradingMaths grading failures 14/20 → 6/20
Reference-guided judgeSolve first, then show that solution as a reference when gradingMaths grading failures 14/20 → 3/20 (70% → 15%)
Fine-tuned judgeTrain a smaller model (Vicuna-13B) on Arena votesPromising preliminary results (Appendix F)

Multi-turn questions

Splitting a two-turn conversation into two separate prompts confused the judge about which earlier answer the follow-up referred to. Showing each full conversation in one prompt, and asking the judge to focus on the second turn, fixed it.

Why it matters today

These fixes are now standard practice: randomise or swap answer order, grade against references where they exist, prefer code-based checks for anything checkable, and keep rubrics specific. See the evaluation lesson.

4 Agreement with humans · original

Everyday picture

If two human judges agree 81% of the time, a machine judge that agrees with humans 85% of the time is doing as well as a human could. The comparison that matters is not “is the judge always right?” but “does it agree with people as often as people agree with each other?”

How agreement is defined

Agreement is the probability that a randomly chosen judge of one type and a randomly chosen judge of the other type give the same verdict on a randomly chosen question. It is reported two ways: S1 counts ties (and position-inconsistent verdicts as ties), so random guessing among A, B and tie would agree 33% of the time; S2 keeps only non-tie votes, so random guessing would agree 50%.

Selected numbers from Zheng et al. (2023), Table 5, MT-Bench first turn. Summarised with attribution.
Pair of judgesS1 (with ties; random = 33%)S2 (no ties; random = 50%)
GPT-4 (pairwise) vs humans66%85%
GPT-4 (single-answer) vs humans60%85%
Human vs human63%81%

Two more findings: when an expert disagreed with GPT-4 and was shown GPT-4's reasoning, they judged it reasonable 75% of the time and changed their own vote 34% of the time. And agreement rose with how different the two models were, from about 70% for close pairs to nearly 100% for very unequal ones.

How much of an agreement rate is real? Cohen's kappa

Raw agreement flatters: two judges who both say “A” most of the time will agree often by chance alone. Cohen's kappa subtracts the agreement expected by chance.

In words: “take the observed agreement, subtract what two judges with the same habits would agree on by luck, and rescale so 1 means perfect agreement and 0 means no better than luck.”

With the numbers: 100 non-tie votes: both say A 40 times, both say B 35 times, judge A while human B 10 times, judge B while human A 15 times. Observed po = 75/100 = 0.75. The judge says A 50% of the time and the human 55%, so pe = 0.50 × 0.55 + 0.50 × 0.45 = 0.50. κ = (0.75 − 0.50)/(1 − 0.50) = 0.50: halfway between luck and perfect.

In Python:

both_A, both_B, judge_A_human_B, judge_B_human_A = 40, 35, 10, 15
n = both_A + both_B + judge_A_human_B + judge_B_human_A
# observed agreement
p_o = (both_A + both_B) / n
# how often the judge says A
p_JA = (both_A + judge_A_human_B) / n
# how often the human says A
p_HA = (both_A + judge_B_human_A) / n
# agreement by luck alone
p_e = p_JA * p_HA + (1 - p_JA) * (1 - p_HA)
p_o, round(p_e, 2)  # → (0.75, 0.5)
# κ
round((p_o - p_e) / (1 - p_e), 2)  # → 0.5

Reading it: the four sliders fill a 2 × 2 table of how often the judge and the human chose A or B. The bars compare raw agreement with kappa. Try making both judges say “A” almost always (raise “both A”, lower everything else except a few disagreements): raw agreement stays high while kappa drops, because most of that agreement is what anyone with the same habit would produce. Kappa is the number to report when you validate a judge; the metrics lesson builds it from scratch.

5 Two kinds of benchmark · original

Everyday picture

A driving test has a written part (knowledge) and a road part (practice). Passing one does not guarantee the other.

Selected rows from Zheng et al. (2023), Table 9. MMLU measures knowledge (5-shot, %); MT-Bench is GPT-4's score out of 10. Summarised with attribution.
ModelFine-tuning dataMMLUMT-Bench
LLaMA-7B (base)none35.22.74
Vicuna-7B (selected)4.8M tokens (about 3K conversations)37.35.95
Vicuna-7B (all)370M tokens47.16.00
GPT-4unknown86.48.99

Reading it: compare the first two rows. A small set of good conversations more than doubled the MT-Bench score (2.74 → 5.95) while barely moving MMLU (35.2 → 37.3). Style and helpfulness can be taught cheaply; knowledge cannot. The paper's recommendation is to report both kinds of benchmark.

6–7 Discussion and conclusion · original

The authors note what they did not cover: the study measures helpfulness and largely ignores safety, honesty and harmlessness, and it folds accuracy, relevance and creativity into one score. They see the same method extending to those dimensions with different prompts. Their conclusion: with its biases addressed, a strong LLM judge is a scalable, explainable approximation of human preference, agreeing with experts at human-to-human levels.

Using it today

PracticeWhyLearn it
Grade with code wherever you can (exact match, schema checks, database state, tests)Cheap, deterministic and bias-free; save the LLM judge for truly open-ended outputevals
Validate the judge on a human-labelled sample; report agreement and kappaA judge you have not checked is an opinion, not a measurementevals
Swap or randomise answer order; call inconsistent verdicts tiesPosition biasevals
Use a specific rubric and, where possible, a reference answerMaths and reasoning grading failures; verbosity biasevals
Re-check agreement whenever the judge model or prompt changesAbsolute scores drift between judgesdeployment

Glossary

Every term with hover guidance on this page, in one place.