Measuring Massive Multitask Language Understanding, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
- The chance explorer lets you set a subject's size and score and see whether the score can be told apart from guessing.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The benchmarks lesson builds the scoring rules, error bars and contamination checks this page keeps pointing at, in plain Python.
Abstract
“We propose a new test to measure a text model's multitask accuracy.”Hendrycks et al. (2020), Abstract. Read the original
Everyday picture
Imagine a single exam paper stapled together from 57 different exams: a school maths quiz, a medical licensing exam, a bar exam, a college physics final, a world religions test. A person who has read widely would pass some sections and fail others. This paper gives that paper to language models, without teaching them anything for it first, and asks how much of what they read during pretraining they can actually use.
What the paper claims
- Most models of 2020 scored close to random chance: 25%, since every question has four options.
- The largest GPT-3 (175 billion parameters) reached 43.9%, almost 20 points above chance, but below expert level on every one of the 57 subjects.
- Knowledge was lopsided: about 70% on its best subject, near chance on others, with calculation-heavy subjects and subjects about human values (law, morality) among the worst.
- Models did not know when they were wrong: their confidence was a poor guide to their accuracy.
Why it matters today
MMLU became a standard yardstick of what a model knows: later papers on this site report it as a matter of course (the QLoRA and MT-Bench companions both show 5-shot MMLU scores). Reading that number well means knowing exactly how it is produced, which is what this page takes apart.
1 Introduction · original
“We design the benchmark to measure knowledge acquired during pretraining by evaluating models exclusively in zero-shot and few-shot settings.”Hendrycks et al. (2020), §1
Everyday picture
Earlier benchmarks were being outgrown within a year: GLUE, then SuperGLUE, reached human-level scores almost as soon as they appeared. The authors argue those tests checked language skills that nearly every child has, while models were reading all of Wikipedia, thousands of books and much of the web. So they built a test of the specialised things people have to study to learn, and they ask it cold: the model sees at most five worked examples in its prompt, and is never trained on the test's format.
Tiny example: one question, asked the paper's way
Every question arrives inside the same frame: a one-line header naming the subject, up to five worked examples, the question with its four options, and the word “Answer:”. The model's job is to continue with one letter. Hover or tap each part.
Hover or tap a part. Start with the header, then follow the prompt down to the prediction.
Reading it: read top to bottom. Everything inside the dashed frame is text the model reads; nothing is trained. The header (yellow) names the subject; the dashed box is where up to five solved examples from the small dev set go, and is empty in the zero-shot setting. The question (pink) and its four options (blue) are followed by “Answer:”, so the most natural next token is a letter. The model then assigns each of the four letters a probability (green); the highest, B, is the prediction, and it is compared with the answer key. The model never writes an explanation, and no parser is needed.
What they found, in one line each
- Few-shot GPT-3 models up to 13 billion parameters sit at chance (about 25%); the 175-billion one reaches 43.9%.
- GPT-3's best subject is almost 70%; several are near chance.
- Its average confidence can be up to 24 points away from its actual accuracy.
Why it matters
The design choice in the quote is the important one. Because the model is never trained on the test, a score measures what was absorbed from ordinary reading, not what was learned from a large pile of practice questions. That is also why, later, the worry about contamination matters so much: if the questions leak into the reading, the test no longer measures what it was built to measure.
2 Related work · original
Everyday picture
A spelling test of “cat” and “dog” cannot tell a ten-year-old from a novelist. Commonsense benchmarks such as HellaSwag, Physical IQA and CosmosQA test what “almost every child has”, and models were already close to human level on them. The paper wants harder, specialised questions.
Two choices the paper argues for
- Few-shot, not fine-tuned. Earlier practice fine-tuned a model on each task's training set. GPT-3 showed that a large model can do many tasks from a handful of examples in the prompt (GPT-3 companion), so a diverse test no longer needs a training set per task, and the model cannot learn the dataset's accidental shortcuts.
- Multiple choice, not free text. Generated text is notoriously hard to grade automatically, and has no standard metric. Multiple choice is scored by accuracy, with no judgment calls.
Why it matters
Both choices bought a simple, cheap, comparable number, and both have a price. Multiple choice measures recognition of the right option, not the ability to explain or apply it. The benchmarks lesson sets this kind beside exact-match maths, unit-tested code and human-preference arenas, each measuring something narrower than its name.
3 A multitask test · original
“Each subject contains 100 test examples at the minimum, which is longer than most exams designed to assess people.”Hendrycks et al. (2020), §3
Everyday picture
Students collected the questions by hand from freely available practice material: practice questions for the Graduate Record Examination and the United States Medical Licensing Examination, questions written for undergraduate courses, and questions for readers of Oxford University Press books. Some subjects come at several levels (“High School Psychology” and “Professional Psychology”, for instance).
Tiny example: where every question goes
The paper collected 15,908 questions and split them three ways. The dev set holds exactly 5 questions per subject, 5 × 57 = 285, and exists only to supply the worked examples in few-shot prompts. The validation set (1,540 questions) is for tuning hyperparameters. The test set (14,079 questions) produces the score.
Hover or tap a box: the four groups at the top, then the three splits.
Reading it: the top row is the 57 subjects sorted into the paper's four groups, with how many subjects each holds (counted from its Table 2). They pour into one pool of four-option questions, which splits three ways. Only the test split (green, thick border) produces the reported score; the dev split's only job is to fill the prompt's worked-example slot in Figure 1, and there is no large training split at all. One small puzzle: 285 + 1,540 + 14,079 = 15,904, four fewer than the 15,908 the paper states.
The math: what does “random chance” look like on 100 questions?
With four options, a model that guesses gets 25% on average. But a subject has only about 100 test questions, so even pure guessing wobbles. The standard error says how much.
In words: “the typical distance between a measured score and the true rate is the square root of the true rate times its complement, divided by the number of questions.”
With the numbers: a guesser has p = 0.25; on a 100-question subject, SE = √(0.25 × 0.75 / 100) = 0.043, so 95% of guessing runs land between 25 − 1.96 × 4.3 = 16.5% and 33.5%. A model that scores 30% on such a subject cannot be told apart from a guesser. The full test is different: with 14,079 questions, GPT-3's 43.9% carries a margin of only about ±0.8 points.
In Python:
import math
# p: a guesser's chance per question; n: questions in one subject
p, n = 0.25, 100
SE = math.sqrt(p * (1 - p) / n)
round(SE, 4) # → 0.0433
# the 95% range of a guesser's score on this subject
round(p - 1.96 * SE, 3), round(p + 1.96 * SE, 3) # → (0.165, 0.335)
# the whole test: 43.9% on 14,079 questions, margin in points
round(100 * 1.96 * math.sqrt(0.439 * 0.561 / 14079), 1) # → 0.8
This calculation is ours, not the paper's; it uses the formula the benchmarks lesson builds as standard_error and confidence_interval.
How good is a person?
People hired through Amazon Mechanical Turk, with no special training, scored 34.5%. For experts, the authors take the 95th-percentile score of real test-takers on the exams the subjects come from (about 87% on the US medical licensing exams behind Professional Medicine) and make an educated guess where no such figure exists, arriving at an estimated expert level of about 89.8%.
Why it matters
The two ends of the scale are fixed by design: 25% is knowing nothing, and about 90% is a well-trained human specialist. Any single subject is a small sample, so per-subject scores near chance are noisy; the whole test, with over fourteen thousand questions, pins the average down to within a point.
3.1 to 3.4 Four groups of subjects · original
Everyday picture
A university prospectus: faculties of humanities, social science and science, and then a long list of professional schools and odd courses that fit nowhere else.
Tiny example: a few subjects from each group
| Group | Some of its tasks | What they ask for |
|---|---|---|
| Humanities (§3.1) | Professional Law, Formal Logic, Moral Scenarios, High School European History | applying rules to messy scenarios; logical fallacies and formal logic; widespread moral intuitions |
| Social science (§3.2) | High School Microeconomics, Econometrics, Security Studies, Professional Psychology | a mix of world knowledge, qualitative and quantitative reasoning |
| STEM (§3.3) | Elementary Mathematics, College Mathematics, Conceptual Physics, Abstract Algebra | procedural problem solving, from arithmetic word problems to GRE-level maths |
| Other (§3.4) | Professional Medicine, Professional Accounting, Global Facts, Marketing | the long tail: business, health, and statistics about the world |
Reading it: the middle column shows the spread of levels inside each group, from elementary school to professional licensing. The right column shows the kind of skill each group leans on: humanities questions often need a rule applied to a detailed scenario (the paper's Figure 2 is a law question of that kind), and STEM questions often need a calculation carried out, which in text means writing maths with LaTeX or with symbols such as * and ^.
Why it matters
Because every question carries a subject label, one run of the test produces 57 scores, not one. That granularity is what lets the paper find blind spots, and it is also what makes per-subject noise (the section above) worth keeping in mind.
Try it: is a subject above chance?
Everyday picture
Toss a coin 10 times and 7 heads is unremarkable; toss it 1,000 times and 700 heads is a scandal. The same score means more on a bigger test.
Try it: set a subject's number of questions and a model's score on it. The band is the 95% range the score could have come from by luck alone. Watch whether it covers the chance line at 25%. Start at 30% on 100 questions, then raise the questions to 1,000.
Reading it: the horizontal line runs from 0% to 100%. The dashed marks are fixed: chance at 25% and the paper's estimated expert level at 89.8%. The dot is the measured score and the shaded band its 95% range, score ± 1.96 standard errors. At 30% on 100 questions the band reaches down past 25%, so the result is compatible with guessing; at 1,000 questions the same 30% is clearly above chance. Scores near 50% have the widest bands, and scores near the top the narrowest.
Why it matters
When a paper says a model is “near random” on a subject, this is the yardstick: a band around the score that still covers 25%. A subject with 100 test questions has a band about ±9 points wide at a score of 30%, so differences of a few points on one subject are usually noise.
4 Experiments · original
Everyday picture
Two kinds of student sit the exam. GPT-3, in four sizes, reads a few worked examples and answers. UnifiedQA was trained beforehand on other question-answering datasets (a different set of practice papers) and sits this one without further practice. Three smaller models, RoBERTa, ALBERT and GPT-2, were fine-tuned on UnifiedQA's multiple-choice questions and MMLU's own dev and validation questions.
4.1 Setup: how a letter is chosen · original
“The model then produces probabilities for the tokens ‘A,’ ‘B,’ ‘C,’ and ‘D,’ and we treat the highest probability option as the prediction.”Hendrycks et al. (2020), §4.1
Everyday picture
Instead of letting the student write an essay, you watch which of four buttons their finger drifts toward, and take the one it drifts toward most.
Tiny example
In Figure 1, after reading the prompt ending in “Answer:”, the model's next-token probabilities for the four letters are (illustratively) A 0.06, B 0.81, C 0.07, D 0.06. The largest is B, so the prediction is B. The key says B: a score of 1 for this question.
The math: the prediction
In words: “the prediction is whichever of the four letters the model thinks is most likely to come next after the prompt.”
With the numbers: P(A | x) = 0.06, P(B | x) = 0.81, P(C | x) = 0.07, P(D | x) = 0.06; the largest is 0.81, at ℓ = B, so ŷ = B.
In Python:
# P(ℓ | x): the model's next-token probability for each letter (illustrative)
P = {"A": 0.06, "B": 0.81, "C": 0.07, "D": 0.06}
# argmax over ℓ
y_hat = max(P, key=P.get)
y_hat # → 'B'
This is one of the two ways the benchmarks lesson scores multiple choice. The other asks how probable each option's full text is (option_logprob, pick_option); comparing single letters sidesteps that method's bias toward short options. Only the four letters are compared, so the model cannot “miss” by writing something else: if it would rather write “The”, that probability is simply ignored.
The math: the score
In words: “accuracy is the share of questions where the predicted letter matches the key.”
With the numbers: five illustrative questions with predictions (B, C, A, D, B) and keys (B, A, A, D, C): matches 1, 0, 1, 1, 0, so acc = 3 / 5 = 0.6.
In Python:
# ŷ_i and y_i for five questions (illustrative)
y_hat = ["B", "C", "A", "D", "B"]
y = ["B", "A", "A", "D", "C"]
N = len(y)
# (1/N) Σ_i 1(ŷ_i = y_i)
sum(1 for a, b in zip(y_hat, y) if a == b) / N # → 0.6
Averaging over subjects: which average?
The paper computes accuracy “across all examples and tasks”, and its Table 1 calls the result an average weighted accuracy. A small check shows the weighting is real: for GPT-3 X-Large, the plain mean of the four group scores (40.8, 50.4, 36.7, 48.8) is 44.2, but the Average column says 43.9. Averaging over questions lets big groups count for more.
In words: “the overall accuracy is each group's accuracy weighted by how many questions the group has.”
With the numbers (illustrative): a 100-question group at 60% and a 300-question group at 30%. Weighted: (100 × 0.6 + 300 × 0.3) / 400 = 150 / 400 = 0.375. The plain mean of the two group scores would say 0.45.
In Python:
# n_g questions and a_g accuracy per group (illustrative)
n_g = [100, 300]
a_g = [0.6, 0.3]
# Σ_g n_g a_g / Σ_g n_g
sum(n * a for n, a in zip(n_g, a_g)) / sum(n_g) # → 0.375
# the plain mean of the group scores, for contrast
round(sum(a_g) / len(a_g), 2) # → 0.45
# Table 1, GPT-3 X-Large: the plain mean of the four columns
round((40.8 + 50.4 + 36.7 + 48.8) / 4, 1) # → 44.2
Zero-shot and few-shot
Zero-shot, the question follows the header directly. Few-shot, up to five dev-set questions with their answers go first. Using the same five fixed examples per subject for every model keeps the comparison fair: the shots are part of the test, not something each lab picks.
Why it matters
Every piece of this recipe (the header's wording, the number of shots, letter probabilities versus full-option probabilities, question-weighted versus subject-weighted averaging) moves the final number. Two MMLU scores are comparable only when they were produced the same way, which is why later reports state “5-shot” beside the score, and why shared evaluation harnesses exist.
4.2 Results: model size and fine-tuning · original
Everyday picture
On most earlier benchmarks, bigger models improved gradually, each size a little better than the last. On this exam, three GPT-3 sizes sit at the guessing line, and only the largest lifts off.
Tiny example
GPT-3 at 2.7, 6.7 and 13 billion parameters scores 25.9%, 24.9% and 26.0% few-shot: all chance. At 175 billion it scores 43.9%. UnifiedQA with 11 billion parameters, smaller than the third GPT-3, scores 48.9% without ever seeing an MMLU question in training.
Hover the chart, or tab to it and use the arrow keys, to read each model's score.
Reading it: the x-axis is the number of parameters on a log scale (each step is ten times more), and the y-axis is average test accuracy. The flat lines are the two ends of the scale: chance at 25% and the estimated expert level at 89.8%. The GPT-3 line (few-shot) lies on the chance line for three sizes, then jumps at 175 billion: this is the shape of the paper's Figure 1(b), where MMLU, unlike commonsense and language benchmarks, shows no gradual climb. The UnifiedQA line starts above chance even at 60 million parameters (29.3%) and climbs through 3 billion (43.7%) to 11 billion (48.9%). Not drawn: the fine-tuned RoBERTa-base (27.9%), ALBERT-xxlarge (27.1%) and GPT-2 (32.4%). Numbers from Table 1 and Appendix A of Hendrycks et al. (2020); sizes are as the paper states them.
What the paper concludes
Size is a key ingredient: only the largest GPT-3 moves beyond chance, and zero-shot it also does better (about 37.7%, against about 25% for the smaller ones). But fine-tuning on other question-answering data helps too: UnifiedQA beats GPT-3 X-Large with about a sixteenth of its parameters. In Appendix A the paper adds that UnifiedQA's small variants beat similar-sized fine-tuned models, and suggests the reason is the larger pretraining dataset behind its T5 backbone.
Why it matters
This was an early, clear example of a capability that looks absent at small scale and appears at large scale, and of how much the surrounding training (here, fine-tuning on other QA formats) can substitute for raw size. It also shows that a benchmark is most useful when models score between chance and the ceiling; the expected_score curves in the benchmarks lesson show why a test sorts models only in that middle range.
4.2 Results: lopsided knowledge · original
“We speculate that is in part because GPT-3 acquires declarative knowledge more readily than procedural knowledge.”Hendrycks et al. (2020), §4.2
Everyday picture
A person who has read every encyclopedia but never done an exercise can tell you what the order of operations is called, and still get 3 + 4 × 2 wrong. Knowing that (declarative knowledge) is different from knowing how (procedural knowledge).
Tiny example
GPT-3 does better on College Medicine (47.4%) and College Mathematics (35.0%) than on Elementary Mathematics (29.9%): harder subjects beat an easier one. The paper checks that GPT-3 can state the PEMDAS rule for the order of operations, and then finds it does not apply it consistently. Nine of its ten lowest-scoring tasks are STEM subjects that stress mathematics or calculation.
Reading it: each pair of bars is one group of subjects, on an axis from 0 to 100%: plain for GPT-3 X-Large (few-shot), striped for UnifiedQA (transfer, no MMLU training). For both models, STEM is the lowest bar and social science the highest, and the gap between them is about 14 to 16 points. Chance is 25% on every bar. Numbers from Table 1 of Hendrycks et al. (2020).
Reading it: five of GPT-3's 57 subject scores, the ones the paper's text names, on the same 0 to 100% axis. The best, US Foreign Policy, is about 69%; the worst, College Chemistry, about 26%, a hair above chance. The middle three show the “unusual order”: the college subjects beat elementary mathematics. None comes close to the expert level of about 90%. The paper's Figure 6 plots all 57, for both models.
Beyond calculation
Procedure is not the only weak spot. Two verbal tasks, Moral Scenarios and Professional Law, are also among the lowest, which the authors single out as worrying: future systems will need a good grasp of what is legal and what is ethical.
Why it matters
An average hides a profile. Two models with the same MMLU score can have very different strengths, so a single number says little about whether a model will do well on your subject. The 57-way breakdown is the useful output; the average is a summary of it.
4.2 Results: calibration · original
Everyday picture
A weather forecaster who says “70% chance of rain” is calibrated if it rains on about 70% of those days. A forecaster who always says 90% and is right half the time is overconfident, and you cannot use their numbers.
Tiny example (illustrative)
Ten answers. For five, the model's top probability was high (0.9, 0.9, 0.95, 0.85, 0.9) and three were right. For the other five it was low (0.4, 0.35, 0.45, 0.4, 0.4) and two were right. Average confidence: 0.65. Accuracy: 5 / 10 = 0.5. The model claims 15 points more than it delivers, all of it in the confident group: 0.9 claimed, 0.6 delivered.
Hover or tap a bar. Blue bars are what the model claimed; green bars are how often it was right.
Reading it: each pair of bars is one group of answers, sorted by how confident the model was. In each pair, the left bar (blue) is the average confidence the model claimed and the right bar (green) is the share it actually got right. A calibrated model's pairs would be equal heights. Here the unsure answers are honest (0.4 claimed, 0.4 right) and the confident ones are not (0.9 claimed, 0.6 right): the pink sliver is the gap, and the formulas below turn gaps like it into one number.
The math: average confidence against accuracy
In words: “the model's confidence on a question is the probability it gave its chosen letter; average that over the questions and compare it with the accuracy.”
With the numbers: the ten confidences sum to 6.5, so conf = 0.65; acc = 0.5; gap = |0.65 − 0.5| = 0.15. The paper measures this gap subject by subject and finds it reaching 24 points for some subjects zero-shot, and 14 points few-shot.
In Python:
# max_ℓ P(ℓ | x_i): the probability of the chosen letter, per answer (illustrative)
conf_i = [0.9, 0.9, 0.95, 0.85, 0.9, 0.4, 0.35, 0.45, 0.4, 0.4]
# 1 if the chosen letter matched the key
right = [1, 1, 1, 0, 0, 1, 0, 1, 0, 0]
N = len(right)
conf, acc = sum(conf_i) / N, sum(right) / N
round(conf, 2), acc # → (0.65, 0.5)
# |conf − acc|
round(abs(conf - acc), 2) # → 0.15
The math: RMS calibration error
The average can hide trouble: overconfidence in one group and underconfidence in another cancel out. The paper also reports the RMS calibration error, which compares confidence and accuracy group by group and squares the gaps so they cannot cancel. The paper names it without writing it out; the formula below is from the paper it cites, Hendrycks, Mazeika and Dietterich (2019), arXiv:1812.04606, Appendix.
In words: “sort the answers into confidence groups, take each group's gap between accuracy and confidence, square it, average the squares with each group weighted by its size, and take the square root.”
With the numbers: two groups of 5, weight 5 / 10 = 0.5 each. Confident group: 0.6 − 0.9 = −0.3, squared 0.09. Unsure group: 0.4 − 0.4 = 0. RMS = √(0.5 × 0.09 + 0.5 × 0) = √0.045 = 0.212. That is larger than the 0.15 average gap, because squaring lets the confident group's big miss dominate. The paper's example: Elementary Mathematics has a zero-shot RMS calibration error of 19.4%.
In Python:
import math
conf_i = [0.9, 0.9, 0.95, 0.85, 0.9, 0.4, 0.35, 0.45, 0.4, 0.4]
right = [1, 1, 1, 0, 0, 1, 0, 1, 0, 0]
N = len(right)
# B_j: the answers in each confidence group (indices)
bins = [range(0, 5), range(5, 10)]
# acc(B_j) and conf(B_j) for each group
acc_b = [sum(right[k] for k in B) / len(B) for B in bins]
conf_b = [round(sum(conf_i[k] for k in B) / len(B), 2) for B in bins]
acc_b, conf_b # → ([0.6, 0.4], [0.9, 0.4])
# √( Σ_j |B_j|/N (acc − conf)² )
round(math.sqrt(sum(len(B) / N * (a - c) ** 2 for B, a, c in zip(bins, acc_b, conf_b))), 3) # → 0.212
Why it matters
A model that is right 44% of the time is much more useful if it can tell you which 44%. The paper found GPT-3's confidence only weakly related to its accuracy zero-shot (a correlation of 0.63 across subjects, rising to 0.81 few-shot). The same concern runs through later work: agreement between samples as a confidence signal in the self-consistency companion, and the losses lesson on how cross-entropy training pushes probabilities toward calibration.
5 Discussion · original
“Learning the entire law exclusively through a small number of practice tests is implausible, so future models must learn more during pretraining.”Hendrycks et al. (2020), §5
Everyday picture
People learn law mostly by reading law, not by memorising thousands of past exam questions; there are not enough past questions to learn it that way. The paper argues models should be tested the same way: read widely, then sit an exam in a format they did not train on.
Tiny example: more reading helps only a little
The authors tried to build a better Professional Law model. RoBERTa-base fine-tuned on about 2,000 extra law questions scored 32.8%. Continuing its pretraining on about 1.6 million legal case summaries first, then fine-tuning, raised that to 36.1%: better, but still far below the expert level.
The three threads of the discussion
- Text only. Much of what people know comes through images, sound and physical interaction. The test is text-only by design; the authors expect benchmarks to follow models into other modalities.
- The Internet as the training set. Dev, validation and test sets exist per task, but no training set. Because the evaluation format differs from anything seen in pretraining, the test avoids annotation artifacts of the kind a model can exploit when training and test data come from the same collection.
- Model limitations. Poor at calculation, poor at modelling human approval and disapproval (law, morality), and below expert level (90%) on every subject. On scaling: the paper cites the scaling-law estimate that ten times more parameters needs about five times more data, and notes that far less is written about esoteric subjects than about everyday life.
Why it matters
“Test on a format the model never trained on” is the core bet of the benchmark. It is also why leaked questions are so damaging: the design assumes the only route to a right answer is knowledge absorbed from reading.
6 Conclusion · original
The paper introduced a test of how well text models learn and apply knowledge met during pretraining, across 57 subjects and many levels. In 2020, models had only just begun to make real progress on it; their knowledge was lopsided, their confidence untrustworthy, and they were weakest on calculation and on morality and law. The authors offer the test as a way to see those shortcomings clearly.
Appendix A: shots, errors and format · original
Everyday picture
Three small checks a careful examiner makes: does showing more solved examples help, what do the confident mistakes look like, and does the exam's layout change the result?
More shots, higher scores
For GPT-3 X-Large, accuracy rises steadily from zero to five worked examples (its Figure 10), but the whole climb is modest: about 37.7% zero-shot against 43.9% five-shot. Most of what the model knows is available without examples; the examples mostly show it the format.
Confident mistakes are often answers to a slightly different question
Asked how many chromosomes all human somatic cells contain, few-shot GPT-3 answered 23 with 97.5% confidence. The answer is 46; 23 is the number of pairs. Many of its other high-confidence mistakes were likewise right answers to nearby questions.
Format sensitivity
GPT-3's accuracy changed little across formatting choices, but UnifiedQA's did: it expects an input shaped like its training data, ending in an end-of-sequence marker </s>, and removing that marker cost several points. The Moral Scenarios format hurt UnifiedQA substantially.
Why it matters
The error analysis is a reminder that “wrong” on a benchmark can mean a near miss or a total miss, and accuracy counts both the same. The format result is the practical one: a harness that formats prompts differently from what a model expects can cost it points that have nothing to do with knowledge.
Appendix B: was the test memorized? · original
“If they memorized the exact question and answer, then they would attain higher accuracy than their true ability.”Hendrycks et al. (2020), Appendix B.2
Everyday picture
A student who has seen the exam before reads its questions fluently, as if recognising them. If fluency went with high marks, you would suspect a leak.
Tiny example: the check
For each task, the authors measure how probable GPT-3 finds the question text itself (its average log-probability, without the answer): a memorised question should look unusually probable. Then they ask whether tasks with more probable questions are also the tasks with higher accuracy. They are not: the correlation across tasks is −0.43 zero-shot and −0.56 few-shot. A correlation near +1 would have been the warning sign; a negative one is the opposite of what memorisation predicts.
Why it matters
It is an indirect check: it needs only the model, not its training data. The authors add two more reassurances (most questions came from PDFs or web pages with answers on separate pages; they would publish a list of sources) and point to the GPT-3 paper's contamination study. The benchmarks lesson builds the direct check, n-gram overlap with the training data (ngram_overlap), and simulates what leaked questions do to a score (simulate_contamination).
What happened next
MMLU became one of the most widely reported benchmarks, and then went through the life cycle the benchmarks lesson describes: useful while scores were spread out, then saturated. By 2024, the authors of MMLU-Pro describe performance on it as having begun to plateau, making differences between models hard to see.
| In the paper | What came later | Learn it |
|---|---|---|
| Four options per question | MMLU-Pro (Wang et al., 2024, arXiv:2406.01574) widened the choice to ten options, removed trivial and noisy questions and added harder reasoning ones; accuracy dropped by 16 to 33 points against MMLU, and scores became less sensitive to prompt wording | saturation |
| Expert level estimated at 89.8% | A re-annotation of 5,700 questions (Gema et al., 2024, arXiv:2406.04127) estimated that 6.49% of MMLU questions contain errors, which caps any honest score below 100% | wrong answer keys |
| Letter probabilities, no explanation | Later evaluations also let models reason before answering; MMLU-Pro found reasoning first helps on its harder questions, in contrast to findings on the original MMLU | chain-of-thought companion |
| An indirect memorisation check | Direct checks by overlap between test items and training data | n-gram overlap |
| One score per subject, one average | Error bars and paired comparisons before calling a gap between two models real | paired bootstrap |
Glossary
Every term with hover guidance on this page, in one place.