An annotated companion · AI Primer

Training Verifiers to Solve Math Word Problems, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is distributed under arXiv's standard non-exclusive licence, which does not grant permission to republish its tables or figures, so none are copied. Its plotted results are redrawn here from scratch: the paper's charts are vector drawings, and the numbers behind every redrawn curve were recovered from the coordinates of the plotted points in its PDF, so they match the original to about a tenth of a percentage point but are not printed anywhere in the paper. The equations are written in this page's own notation for methods the paper describes in words. Numbers marked illustrative are made up for teaching. Read the original alongside: every section links to it.

How to read this page

  • Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
  • The lure explorer lets you change how often a model is right and how often it writes a wrong solution the verifier loves, and shows why sampling more can make things worse.

Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. This paper came a few months before chain-of-thought prompting and self-consistency, and its dataset is the one both of them report on. Its sequel, Let's Verify Step by Step, asks whether a verifier should judge every step instead of the final answer. The reasoning models lesson builds verifiers and best-of-n from scratch.

Abstract

“At test time, we generate many candidate solutions and select the one ranked highest by the verifier.”Cobbe et al. (2021), Abstract. Read the original

Everyday picture

A student who writes one answer and hands it in gets graded on that answer. A student who drafts a hundred answers and has a sharp-eyed friend pick the most convincing one does much better, as long as the friend can tell a sound solution from a plausible-looking wrong one. This paper trains the friend: a second model, the verifier, whose only job is to read a candidate solution and say how likely it is to be right.

What the paper claims

  • A dataset. GSM8K: 8.5K grade-school maths word problems with worked solutions in plain English, hard for 2021's largest models though every problem needs only arithmetic.
  • A method. Sample many solutions from a fine-tuned model, score each with a trained verifier, keep the top one (best-of-n).
  • A result. A 6-billion-parameter model with a verifier slightly beats a 175-billion-parameter model fine-tuned the ordinary way: about the gain of a 30× larger model. And verification keeps improving with more training data, faster than fine-tuning does.

Why it matters today

GSM8K became one of the most reported benchmarks in language modelling. The verifier became the outcome reward model, and “sample many, keep what a checker likes” became a standard way to spend test-time compute.

1 Introduction · original

“When generating a solution, autoregressive models have no mechanism to correct their own errors.”Cobbe et al. (2021), §1

Everyday picture

Writing a solution in ink, one word at a time, with no eraser. One wrong number in line two and every line after it builds on the mistake. A language model writing a solution is in exactly that position: autoregressive generation appends each token and never goes back.

Tiny example: why one slip is fatal

If each of a solution's steps is right with chance 0.9, a 2-step problem is fully right with chance 0.9 × 0.9 = 0.81, and an 8-step problem with chance 0.98 = 0.43 (illustrative numbers). GSM8K problems take 2 to 8 steps, so a model that is “usually right per step” still fails the longer problems more often than not. The reasoning lesson's steps_all_right is this formula.

The two bets the paper makes

  • Optionality. A generator that is right 30% of the time is usually right at least once in 100 tries. Something only has to find that one.
  • Checking is easier than doing. Judging whether a finished solution holds together is, in general, a simpler task than producing one. So a verifier can add accuracy that the generator alone lacks.

Why it matters

The paper's motivating worry is cost: extrapolating 2021's trends, simply growing models would need an “exorbitant” number of parameters to do well on harder maths. It looks for a method with a better scaling law, and finds it in checking rather than generating. Shen et al. (2021) was concurrent work on the same idea.

2 The GSM8K dataset · original

“A bright middle school student should be able to solve every problem.”Cobbe et al. (2021), §2

Everyday picture

A good test for a reasoner uses easy ingredients in unfamiliar combinations. If every question were the same template with new numbers, a model could memorise the template. If questions needed university maths, failure would tell you nothing about reasoning. GSM8K sits in between: school arithmetic, written fresh by people for every problem.

Tiny example: the solution format

Every problem comes with a worked solution in plain English, one line per step, and the final answer after it. An illustrative problem in the same style: A baker makes 3 trays of 12 rolls and sells all but 5. How many rolls does she sell?

She makes 3 * 12 = <<3*12=36>>36 rolls.
She sells 36 - 5 = <<36-5=31>>31 rolls.
#### 31

The parts in double angle brackets are calculator annotations, added automatically after the fact (the next section explains them). The line starting with #### holds the final answer, the only thing a grader compares.

The four design principles

  • High quality. Problems were written by people, not scraped. After every problem was re-solved by a different worker and disagreements fixed or dropped, a final check on a subset found 1.7% still disputed: the paper's estimate of problems with breaking errors (Appendix A).
  • High diversity. Writers were told not to reuse settings or templates, and pairwise similarity scores between problems were used to catch it. That makes held-out accuracy meaningful.
  • Moderate difficulty. No concept beyond early algebra, 2 to 8 steps, mostly solvable without naming a variable. Hard enough to challenge the largest models, easy enough to measure progress. (The harder MATH dataset was, in 2021, too hard to measure progress on.)
  • Natural-language solutions. Written in each author's own words rather than as bare equations, because the paper expects them to shed light on a model's “internal monologue”.

The split is 7.5K training problems and 1K test problems. (The paper rounds; the GSM8K glossary entry gives the released test set's exact size.)

A small check you can do yourself

The paper's Figure 1 shows three problems. In the third, 6 people at a party drink 3 × 3 = 9, 2 × 4 = 8 and 1 × 5 = 5 sodas out of 36. The printed solution totals them as 5 + 9 + 8 + 3 = 25, adding the 3 people a second time, and answers 11 sodas left; the arithmetic of its own lines gives 22 drunk and 14 left. Human-written labels are not perfect, which is exactly what the 1.7% estimate says, and it matters later: a verifier learns from labels like these.

Why it matters today

GSM8K became the standard first test of multi-step reasoning, used by the chain-of-thought and self-consistency papers within months. Its format (steps, then a final answer that can be checked exactly) is what makes automatic grading, and later verifiable rewards, possible. Modern models now score so high on it that it is close to saturated; the benchmarks lesson covers how a benchmark wears out.

3 Related work · original

Everyday picture

Before this paper, maths word problem datasets were either small, templated, answer-only, or noisy, and most methods built special machinery (equation trees, custom encoders) for the task. The paper's approach is the opposite: plain language in, plain language out, a general model, and a second general model to check it.

What it compares itself with

  • Datasets. Dolphin18K (answers or equations only), AQuA-RAT (100K problems but heavily templated), MathQA (a cleaned subset of AQuA-RAT that still has quality issues in about 30% of items), Ape210K (Chinese, no natural-language solutions), ASDiv (2.3K problems, closest in spirit, fewer steps), and MATH (harder).
  • Methods. Recurrent and encoder-decoder solvers, extra pretraining on maths, and, closest to this paper, methods that train a model to select among many completions: sample-and-rank for storytelling (Nichols et al., 2020) and generate-and-rank for maths (Shen et al., 2021).

Why it matters

Against generate-and-rank, the paper claims three differences: natural-language solutions rather than expressions, evidence that verifiers scale better with data, and separate generator and verifier networks, so the generator can be kept from overfitting while the verifier trains.

4 Methods · original

Everyday picture

Two ways to get better answers from the same student. Finetuning: drill them on worked examples, then take their single most careful answer. Verification: drill them lightly, have them dash off a hundred answers, and let a trained marker pick one.

Setup

Both methods start from GPT-3 models, mostly the 6B and 175B sizes (the 6B is “significantly more convenient for research”). Finetuning is the baseline: fine-tune on the training solutions with the ordinary next-token cross-entropy loss, then sample one solution at temperature 0 (always the most likely token, greedy decoding) and check its final answer. Verification samples many solutions at a higher temperature and lets a verifier choose.

Why it matters

Because the two methods share the same starting models and data, any gap between them comes from the choosing, not from a better base model.

Calculator annotations · original (Appendix C)

Everyday picture

A student who reasons well but adds badly is allowed a calculator. They still decide what to compute; the calculator only does the sum. The paper does this for its models, because even large models often got the arithmetic wrong.

Tiny example

The model writes Her sister gave her 20 + 10 = <<20+10=. The moment it writes “=” inside the brackets, sampling stops and a calculator evaluates 20+10, writes 30>> into the text, and hands control back. The model then continues from the correct number.

text so farwriternext … 20 + 10 = <<20+10 model = “=” inside the brackets: the calculator takes over … = <<20+10= eval 30>> the result is written; the model resumes … <<20+10=30>> model books During training the brackets are ordinary tokens; only at test time does the calculator step in.

Hover or tap a box. Read the rows top to bottom: the model writes, the calculator interrupts, the model resumes.

The calculator sampling procedure (the paper's Figure 9), redrawn with its example sentence. Based on Cobbe et al. (2021), Appendix C.

Reading it: each row is one moment of generation: on the left the text so far, in the middle who writes the next piece, on the right what gets written. In the first row the model is writing as usual and produces “=” inside the double brackets. That token is the trigger: in the second row the writer is no longer the model but a calculator (the green boxes), which evaluates the expression in the brackets and writes the result and the closing brackets. In the third row the model is back in charge and continues the sentence (“books”) from a correct result. The two “model” boxes light up together because they are the same network.

How it is built

  • The annotations were generated automatically, by hard-coded rules plus a fine-tuned model. The paper judges them very unlikely to be wrong, though they often skip lines that could have been annotated.
  • The calculator is Python's eval on the expression. If it errors or times out, the annotation is skipped and the model samples as usual.
  • The paper discloses that its calculator had minor bugs in every reported run, so results slightly understate what is possible: fixing it adds about 1% to verification on the full training set.

Why it matters today

This is an early, minimal form of tool calling: the model learns when to call a tool from the format of its training data, and the system intercepts that call. Modern models call calculators, code interpreters and search the same way, through structured tool calls; the tools lesson builds that loop.

4.1 Finetuning · original

“This is to be expected: as the model repeatedly encounters the same data, it becomes increasingly uncalibrated and overconfident in its predictions.”Cobbe et al. (2021), §4.1

Everyday picture

More drilling makes a student's best single answer a little better, but also makes every answer look the same. Ask them a hundred times and you hear the same attempt a hundred times. For verification, variety is the raw material, so the paper has to measure both.

Tiny example: test@1 and test@3

The paper writes test@N for the share of test problems solved at least once in N guesses. Four problems, three guesses each (1 = right, illustrative): problem A ✗✓✗, B ✗✗✗, C ✓✓✗, D ✗✗✓. With only the first guess, one problem (C) is solved: test@1 = 25%. With all three, A, C and D are each solved at least once: test@3 = 75%.

The math: test@N

In words: “for each problem, ask whether any of its first N guesses was right; test@N is the percentage of problems where the answer is yes.”

With the numbers: P = 4. With N = 1 the maxima are 0, 0, 1, 0, which sum to 1, and 100/4 × 1 = 25. With N = 3 they are 1, 0, 1, 1, summing to 3: 75. If each guess is right with the same chance p, independently, test@N is expected to be 1 − (1 − p)N, the pass@n formula.

In Python:

# c[j][i]: was guess i on problem j right? (illustrative)
c = [[0, 1, 0], [0, 0, 0], [1, 1, 0], [0, 0, 1]]
P = len(c)
# test@1: only the first guess counts
100 * sum(max(row[:1]) for row in c) / P  # → 25.0
# test@3: any of the three guesses
100 * sum(max(row[:3]) for row in c) / P  # → 75.0

Figure 2: bigger models, more data

The paper fine-tunes four GPT-3 sizes on training sets from 500 problems to the full 7.5K, for 20 epochs each, and measures test@1.

Hover the chart, or tab to it and use the arrow keys, to read each model's solve rate.

Reading it: the x-axis is the number of training problems on a doubling (log) scale, ending at the full training set; the y-axis is the percentage of test problems solved with one greedy sample. Each line is one model size, and the lines never cross: bigger is better at every data size. Follow the right-hand edge: with all the data, 3B solves 17.3%, 6B 20.6%, 12B 24.9% and 175B 34.3%. Even the largest model, trained on everything, fails two problems in three. Redrawn from the paper's Figure 2 (left panel); values recovered from the plotted points.

The math: how big would a model need to be?

The paper assumes a log-linear trend: each tenfold increase in parameters adds the same number of points. Fit a straight line through the four right-hand points against the logarithm of the model size, and ask where it reaches 80%.

In words: “the solve rate is a starting value plus a fixed number of points for every tenfold increase in model size.”

With the numbers: a least-squares fit through (3B, 17.3), (6B, 20.6), (12B, 24.9) and (175B, 34.3) gives b ≈ 9.5 points per tenfold and a ≈ −72.0. Setting r = 80 gives log10 M = (80 + 72.0) / 9.5 ≈ 16.0: a model of about 1016 parameters, which is exactly the paper's naive extrapolation. That is about 57,000 times GPT-3.

In Python:

import math
# right-hand points of Figure 2: model size, % solved with all the data
M = [3e9, 6e9, 12e9, 175e9]
r = [17.33, 20.59, 24.94, 34.34]
x = [math.log10(m) for m in M]
x_bar, r_bar = sum(x) / len(x), sum(r) / len(r)
# least squares: b = Σ (x − x̄)(r − r̄) / Σ (x − x̄)²
b = sum((xi - x_bar) * (ri - r_bar) for xi, ri in zip(x, r)) / sum((xi - x_bar) ** 2 for xi in x)
round(b, 2)  # → 9.49
a = r_bar - b * x_bar
round(a, 1)  # → -72.0
# solve 80 = a + b · log10(M) for log10(M)
round((80 - a) / b, 1)  # → 16.0
# how many GPT-3s is that?
round(10 ** ((80 - a) / b) / 175e9, -3)  # → 59000.0

(The ratio lands at about 59,000 unrounded; 1016 / 175 billion is 57,000. Either way, not a model anyone could train.) Along the data axis the curve is not log-linear, so the paper only guesses that the 175B model would need at least a hundred times more training data to reach 80%.

Figure 3: better single answers, less variety

Show:

Hover the chart, or tab to it and use the arrow keys, to read the solve rate after each epoch.

Reading it: the x-axis is how many passes (epochs) over the full training set the 6B model has made; the y-axis is the solve rate, and each button switches the measure. test@1 jumps in the first two epochs and then creeps upward, from about 18.5% to about 21.7% at epoch 50. test@100 does the opposite after a brief rise: it peaks at 83.9% at epoch 3 and slides to about 70.6% by epoch 50. Training longer makes the best single guess slightly better and the spread of guesses much worse, because the model grows overconfident and keeps sampling the same few solutions. Redrawn from the paper's Figure 3; values recovered from the plotted points. (The text describes the curves as covering 100 epochs; the plotted axis stops at 50.)

Two practical findings

  • Pick the generator for coverage, not for its best guess. Because test@100 peaks early, the generator that feeds the verifier is trained for only 2 epochs.
  • Let it show its work. A 6B model fine-tuned to output the final answer directly, with no steps, drops from 20.6% to 5.2%. The written steps are doing work, the same finding as chain-of-thought prompting a few months later.

Why it matters

“Accuracy of one answer” and “chance that some answer is right” are different targets, and optimising the first can wreck the second. Anything that samples many times (best-of-n, self-consistency, reinforcement learning from a group of attempts) needs the second. Calibration is the thread: an overconfident model puts nearly all its probability on one path.

4.2 Verification · original

“Training solutions are labeled as correct or incorrect based solely on whether they reach the correct final answer.”Cobbe et al. (2021), §4.2

Everyday picture

To train a marker, you don't need a teacher to annotate every mistake. Collect a pile of real student answers, check each final number against the answer key, and let the marker learn from thousands of “this one was right, this one wasn't”. The marker has to work out on its own what right solutions look like.

Tiny example: best of four

Four sampled solutions to one problem, with the verifier's scores (illustrative): 0.12, 0.81, 0.47 and 0.66. Best-of-4 returns the second, the one the verifier thinks most likely correct, and only then is its final answer checked against the key.

1 · Train the generator question + solution generator, 2 epochs 2 · Generate and label 100 solutions per problem question generator solution 1 solution 2 … solution 100 ✓ ✗ ✓ 3 · Train the verifier, 1 epoch question + solution + label verifier 4 · Test time new question generator 100 samples, T = 0.7 verifier scores each return the top one The generator never learns from the verifier.

Hover or tap a box. Follow the numbered rows: steps 1 to 3 build the verifier, step 4 uses it.

The verification training pipeline (the paper's Figure 4) redrawn, with the test-time step added as row 4. Based on Cobbe et al. (2021), §4.2.

Reading it: read the four rows top to bottom. Row 1 fine-tunes the generator briefly, so it keeps the variety Figure 3 showed it loses with longer training. Row 2 uses that generator to write 100 solutions for every training problem, and labels each one ✓ or ✗ by comparing its final answer with the key: no person reads the reasoning. Row 3 trains the verifier on the resulting pile of (question, solution, label) triples. Row 4 is the payoff: for a new question, sample 100 solutions and return the one the verifier scores highest. The three purple “generator” boxes are one model and light up together; so do the two “verifier” boxes.

The math: choosing the best of N

In words: “score every sampled solution with the verifier, given the question, and return the solution with the highest score.”

With the numbers: N = 4, scores 0.12, 0.81, 0.47, 0.66. The largest is 0.81, belonging to solution 2, so ŝ = s2.

In Python:

# V(q, s_i) for N = 4 sampled solutions (illustrative)
V = [0.12, 0.81, 0.47, 0.66]
# argmax over i; Python counts from 0, the page from 1
i_hat = max(range(len(V)), key=lambda i: V[i])
i_hat + 1  # → 2
V[i_hat]  # → 0.81

The math: what the verifier is trained to predict

A verifier is itself a language model. After every token of a solution it outputs a number vt, and training pulls each of those numbers towards the solution's label y: 1 if the final answer was right, 0 if not. The paper uses mean squared error for this (Appendix B; cross-entropy worked about as well). It also keeps training the verifier on the ordinary language-modelling loss, and simply adds the two.

In words: “the verifier's loss is its usual next-token loss plus, averaged over the solution's tokens, the squared distance between its running guess and the solution's true label.”

With the numbers (illustrative): a 4-token solution that reached the right answer (y = 1), with verifier outputs 0.3, 0.5, 0.6, 0.8. The squared errors are 0.49, 0.25, 0.16, 0.04; their mean is 0.235. Had the answer been wrong (y = 0), the same outputs would cost 0.335. A solution-level verifier would be graded only on its last output: (0.8 − 1)2 = 0.04. With an illustrative language-modelling loss of 1.2, the total is 1.2 + 0.235 = 1.435.

In Python:

# v_t after each of T = 4 solution tokens (illustrative)
v = [0.3, 0.5, 0.6, 0.8]
T = len(v)
# y = 1: this solution's final answer matched the key
y = 1
[round((v_t - y) ** 2, 2) for v_t in v]  # → [0.49, 0.25, 0.16, 0.04]
round(sum((v_t - y) ** 2 for v_t in v) / T, 3)  # → 0.235
# the same outputs if the answer had been wrong
round(sum((v_t - 0) ** 2 for v_t in v) / T, 3)  # → 0.335
# a solution-level verifier is graded on the last token only
round((v[-1] - y) ** 2, 2)  # → 0.04
# add the language-modelling loss, unweighted (illustrative 1.2)
L_LM = 1.2
round(L_LM + sum((v_t - y) ** 2 for v_t in v) / T, 3)  # → 1.435

Only the solution's tokens count; the question's are masked out of both losses (Appendix E), the same move as an response_mask in supervised fine-tuning. Because the verifier data has 100 solutions per problem, mixing it half and half with the original training solutions means the paper effectively shows each original solution 100 times.

Figure 5: the main result

Model size:

Hover the chart, or tab to it and use the arrow keys, to compare the two methods at each data size.

Reading it: the x-axis is the number of training problems (doubling scale), the y-axis the percentage of test problems solved; one line is ordinary finetuning (one greedy answer), the other verification (best of 100 samples). On the left of each chart verification loses: with 500 problems, the 6B verifier solves 1.9% against finetuning's 7.5%. For 6B the lines cross between 2,000 and 4,000 problems, and from there verification pulls away: on the full set, 38.8% against 20.6% for 6B, and 55.9% against 34.3% for 175B. Switch to 175B and notice that its verifier crosses over earlier, between 1,000 and 2,000 problems: the paper's observation that larger verifiers “take off” sooner. Redrawn from the paper's Figure 5; values recovered from the plotted points. (The 175B verification line is one run; the others average three.)

Why small datasets hurt verifiers

With few problems, a verifier can score well by memorising which final answers are right for the training questions, rather than learning what sound reasoning looks like; the paper calls this “pressure to overfit to the correct answer”. The labels are also imperfect: a solution that reaches the right number through wrong reasoning is labelled ✓, a false positive. With enough varied problems, the general pattern wins.

Why it matters

This is the recipe of an outcome reward model: labels from final answers alone, used to rank samples. The reasoning lesson's verifier_experiment compares best-of-n against a single sample and a vote, and outcome_reward is the labelling rule in one line.

4.3 Verification ablations · original

“This suggests that the verifier may often be relying on relatively coarse heuristics to discriminate between solutions from a given generator, rather than attempting a more thorough form of verification.”Cobbe et al. (2021), §4.3

Everyday picture

A marker who only glances at the boxed answer learns “which answers are usually right”. A marker who has to say, after every line, how things are going learns to read the working. The paper tests which one generalises, and two other choices, by ablation: change one thing, measure again.

Tiny example: token-level against solution-level

The same 4-token solution from the loss example: a token-level verifier is trained on all four outputs, a solution-level verifier on the last one only. At test time both are read at the final token. The token-level one is, in effect, a value function: after each token, its guess at how the solution will end.

The math: the scalar head

How does a language model output a score? The paper reserves one special token in the vocabulary. Its logit is shifted and scaled by two learned numbers, and the result is the verifier's prediction; every other logit keeps doing next-token prediction (Appendix E).

In words: “the prediction at token t is the special token's logit, multiplied by one learned gain and shifted by one learned bias.”

With the numbers (illustrative): a logit of 1.5 with gain 0.4 and bias 0.2 gives 0.4 × 1.5 + 0.2 = 0.8, the last output in the loss example.

In Python:

# z_t: the special token's logit at token t; g, b: the two learned numbers (illustrative)
z_t, g, b = 1.5, 0.4, 0.2
round(g * z_t + b, 2)  # → 0.8

Two parameters on top of a full language model: nearly all of the verifier's judgement lives in the network it shares with language modelling. The paper also found it slightly better to start the verifier from the generator's weights than from the plain pretrained model.

Figure 6: three ablations

Ablation:

Reading it: in (a) and (b) the x-axis is progress through the verifier's single training epoch and the y-axis the 6B verifier's test solve rate. In (a) the solution-level verifier starts faster, peaks near 35% halfway through, then declines to about 30%: overfitting. The token-level verifier starts slower, because predicting after every token is a harder and noisier task, but keeps climbing to about 39%. In (b), dropping the language-modelling loss costs about 3 points at the end (35.4% against 38.8%). In (c) the x-axis is training problems; a 175B generator with a 6B verifier reaches 52.2%, while a 6B generator with a 175B verifier reaches only 42.1%. A small verifier checking a strong generator beats a strong verifier checking a weak one. Redrawn from the paper's Figure 6; values recovered from the plotted points.

Why it matters

The paper's own reading of (c) is modest: the verifier may be leaning on coarse cues rather than truly checking. That caution is worth keeping. A learned verifier is a pattern-matcher trained on labels, and the next section shows what happens when you search hard against one. Judging each step, the idea behind (a), is what Let's Verify Step by Step scales up with human labels on every step.

5 Additional experiments · original

Everyday picture

Once you have a marker, two dials remain. How many drafts do you ask for? And how do you keep both the student and the marker from memorising the practice sheets?

What follows

§5.1 turns the first dial (completions per problem, and voting among the top-ranked ones); §5.2 turns the second (dropout).

Why it matters

Both are still the questions anyone deploying best-of-n asks: how much sampling is worth paying for, and how to stop a learned checker from being brittle.

5.1 Test-time compute · original

“This suggests that the benefits of search are eventually outweighed by the risk of finding adversarial solutions that fool the verifier.”Cobbe et al. (2021), §5.1

Everyday picture

A hundred drafts give the marker plenty of chances to find a right one. Three thousand drafts also include a few wrong ones that happen to look exactly like what the marker likes. Keep searching and eventually you find one of those.

Tiny example: voting among the top-ranked

Six sampled solutions to a problem whose answer is 18, ranked by verifier score (illustrative): 0.91 → 26, 0.88 → 18, 0.85 → 18, 0.62 → 26, 0.55 → 26, 0.30 → 11. The top one alone says 26, wrong. Let the top 3 vote and 18 wins, two votes to one. Let all 6 vote and 26 wins again, three votes to two. Too few voters trust one lucky high score; too many let low-ranked solutions back in.

The math: a vote among the top k

In words: “keep the k solutions the verifier ranks highest, count how many of them reach each final answer, and return the answer with the most.”

With the numbers: with k = 1 the only count is 26: 1, so â = 26. With k = 3 the counts are 18: 2 and 26: 1, so â = 18. With k = 6 they are 11: 1, 18: 2, 26: 3, so â = 26.

In Python:

# (verifier score, final answer), already sorted best first (illustrative)
ranked = [(0.91, 26), (0.88, 18), (0.85, 18), (0.62, 26), (0.55, 26), (0.30, 11)]
def vote(k):
    top = [a_i for _, a_i in ranked[:k]]
    # Σ over the top k of 1(a_i = a), for every answer a
    counts = {a: sum(1 for a_i in top if a_i == a) for a in sorted(set(top))}
    return counts, max(counts, key=counts.get)
vote(1)  # → ({26: 1}, 26)
vote(3)  # → ({18: 2, 26: 1}, 18)
vote(6)  # → ({11: 1, 18: 2, 26: 3}, 26)

Figure 7: how many samples, how many voters

Hover the chart, or tab to it and use the arrow keys, to read the solve rate at each number of completions.

Reading it: the x-axis doubles the number of completions the 6B verifier chooses from, from 25 to 3,200 (the verifier was trained on 100 per problem throughout); the y-axis is the test solve rate. The line rises from 34.6% at 25 completions to a peak of 39.6% at 400, then falls: 37.2% at 3,200. More candidates stopped helping and started hurting. The paper settles on 100 as the default, at 38.4%, since that captures most of the gain for a quarter of the cost of 400. Redrawn from the paper's Figure 7(a); values recovered from the plotted points.

Completions:

Reading it: now the verifier ranks a fixed pile of completions, and the x-axis (log scale) is how many of the top-ranked ones get a vote, from 1 (plain best-of-n) to 100. With 100 completions the best is to let the top 3 to 5 vote (39.1%, up from 38.3% for the top one alone); letting all 100 vote falls to 28.4%. Switch to 3,200 completions: the best number of voters grows to about 30 (43.8%), because a bigger pile has more good solutions near the top. Voting also rescues the large piles that hurt plain best-of-n: 3,200 completions with 30 voters beats every setting in the chart above. Redrawn from the paper's Figure 7(b); values recovered from the plotted points.

Why it matters

Voting among verifier favourites combines the two selection rules that self-consistency and this paper each use alone: agreement protects against one lure, the verifier filters out the obvious junk. The follow-up paper, Let's Verify Step by Step, tries a close cousin: votes weighted by the reward model's score.

Try it: why more samples can hurt

Everyday picture

Picture three kinds of solution in the pile. Right ones. Ordinary wrong ones, which the verifier sees through. And rare lures: wrong solutions the verifier scores higher than anything, the “adversarial solutions” of §5.1. Best-of-N returns a right answer only if the pile holds at least one right solution and not a single lure.

Tiny example

Say 30% of samples are right and 1 in 500 is a lure (illustrative). With 10 samples, the chance of at least one right solution and no lure is 0.99810 − 0.69810 = 0.953. With 400 samples it is only 0.449: a right solution is almost certain, but so is a lure.

The math: a toy model of search against a verifier

In words: “the chance best-of-N is right is the chance that no sample is a lure, minus the chance that no sample is a lure and none is right either.”

With the numbers (illustrative): p = 0.3, q = 0.002. N = 1 gives 0.998 − 0.698 = 0.3: with one sample the verifier has no choice. N = 10 gives 0.953, N = 100 gives 0.819, N = 400 gives 0.449. The peak, where adding a sample stops helping, is near N = 15. Without lures (q = 0) the formula becomes 1 − (1 − p)N, the pass@n ceiling, and never falls.

In Python:

import math
# p: share of samples that are right; q: share that are lures (illustrative)
p, q = 0.3, 0.002
def P_right(N):
    # no lure at all, minus: no lure and no right solution
    return (1 - q) ** N - (1 - p - q) ** N
[round(P_right(N), 3) for N in (1, 10, 100, 400)]  # → [0.3, 0.953, 0.819, 0.449]
# the peak: set the slope in N to zero and solve
a, b = 1 - q, 1 - p - q
round(math.log(math.log(b) / math.log(a)) / math.log(a / b), 1)  # → 14.5
# the same p with no lures: the pass@N ceiling
round(1 - (1 - p) ** 400, 3)  # → 1.0

Try it: move the sliders. Watch where the solid line peaks and how far below the dashed ceiling it ends. Then make lures rarer (drag the second slider right) and watch the peak move right and the fall get gentler.

Hover the chart, or tab to it and use the arrow keys, to read the chance at each N.

Reading it: the x-axis is the number of samples N on a log scale, from 1 to 3,200, like Figure 7(a); the y-axis is the chance that best-of-N returns a right solution. The dashed line is pass@N: the chance that any sample is right, which only rises. The solid line is best-of-N with lures in the pile. It follows the ceiling at first, then turns over once lures become likely, and the gap between the two lines is what the verifier's blind spots cost. The real verifier rises more slowly than this toy, because it also sometimes ranks ordinary wrong solutions first, but the turn-over has the same cause. Every number here is from the formula, with illustrative settings.

Why it matters today

This is Goodhart's law in miniature: a learned score used as a target gets exploited, here by nothing more than sampling. The same thing happens when a policy is trained against a learned reward (reward hacking), which is why modern reasoning training prefers verifiable rewards (a program that checks the answer) wherever one exists. The alignment lesson's goodhart_curve draws the same rise and fall.

5.2 Regularization · original

Everyday picture

A study group where, on any given day, a random fifth of the members stay home. Nobody can rely on one particular friend to know the answer, so everyone learns a bit of everything. That is dropout, and it fights memorisation.

Tiny example

With 20% dropout, a layer's output of 10 numbers has on average 2 of them zeroed on each training step, a different 2 each time. The paper applies it along the residual paths of every layer (residual dropout, from the original transformer). Because GPT-3 was pretrained without dropout, the paper first continues pretraining with dropout switched on, so finetuning does not also have to absorb that change.

What 20% dropout changed, 6B models, from the paper's Figure 8 (values recovered from its plotted points).
SettingNo dropoutDropout 0.2
Finetuning, full training set (test@1)20.6%25.7%
Solution-level verifier, end of its epoch29.9%38.2%
Token-level verifier, mean of the last fifth of training40.6%41.4%

Reading it: each row compares the same model with and without dropout. Finetuning gains 5 points. The solution-level verifier gains the most, because dropout cures the overfitting Figure 6(a) showed: with it, the solution-level verifier ends level with the token-level one. The token-level verifier, already resistant to overfitting, gains about a point. (That run used 4× larger batches and 300 completions per problem, so its level is not comparable with the rest of the page.)

Why it matters

Every run in the paper is one epoch or a few over a small dataset with a huge model: the setting where overfitting bites. The regularization lesson's dropout builds the mechanism; large pretraining runs today often skip it, since they rarely see a token twice.

6 Conclusion · original

“On the full dataset, 6B verification slightly outperforms a finetuned 175B model, thereby offering a boost approximately equivalent to a 30x model size increase.”Cobbe et al. (2021), §6

Everyday picture

A small team with a good reviewer outperforms a much bigger team without one.

The math: where “30×” comes from

In words: “the boost is quoted as the size of the larger model divided by the size of the smaller one that matches it once verification is added.”

With the numbers: on the full training set, 6B with verification solves 38.8% and 175B with finetuning 34.3% (Figure 5). 175 / 6 = 29.2, rounded to 30.

In Python:

M_big, M_small = 175, 6
round(M_big / M_small, 1)  # → 29.2
# the two solve rates being compared, from Figure 5 (% of test problems)
verified_6B, finetuned_175B = 38.76, 34.33
round(verified_6B - finetuned_175B, 1)  # → 4.4

Why it matters

The claim is about scaling: verification's line in Figure 5 is steeper, so the gap should widen with more data and harder problems. The paper expected verification to carry over to harder maths, and the next generation of work did exactly that, with finer-grained checking (Let's Verify Step by Step) and with the check turned into a training signal (DeepSeekMath and GRPO).

Appendix F: watching a verifier read · original

Everyday picture

A marker who says out loud, line by line, “fine… fine… hmm… no”. Because a token-level verifier makes a prediction after every token, you can colour each token by it and see where its confidence drops.

Tiny example: the paper's fourth row

The problem: Howard spends 8 dollars at the arcade on Monday, twice as much on Tuesday, and 4 times Tuesday's amount on Wednesday. He started with 100 dollars; how much is left? The right answer: 8 + 16 + 64 = 88 dollars spent, 12 left.

He spent $8 on Monday and $8*2 = $16on Tuesday. He spent $16 on Tuesday and $16*4 = $64on Wednesday. He has $100 and spent $64, so he has100 − 64 = 36 dollars left. #### 36 verifier: wrong actually: wrong green: high score · amber: falling · pink: low

Hover or tap each line to hear the verifier's running verdict.

One row of the paper's Figure 13, redrawn: a 175B solution scored token by token by a 175B token-level verifier. The colours follow the paper's description and shading, simplified to one colour per line. Based on Cobbe et al. (2021), Appendix F.

Reading it: read down the solution. The first line is correct and the verifier is confident (green). The second is also correct arithmetic, but the paper describes the verifier growing gradually less confident as the solution goes on (amber). The third line says Howard spent $64 in total, forgetting Monday and Tuesday; the verifier's score collapses (pink) and it ends confident the solution is wrong, which it is. The bottom boxes compare the verifier's verdict with the truth: they agree.

The other rows

  • True positive. On a correct solution the verifier starts unsure and grows more confident line by line; the paper attributes this to training on many wrong samples.
  • Two false negatives. One correct solution is rejected, perhaps because the problem says both “4 times” and “4 potatoes”. Another reaches the right final answer through faulty reasoning, and the verifier rightly scores it low: its label says correct, its reasoning says otherwise.
  • False positive. A solution subtracts $400 from the wrong jewel's price and the verifier misses it. The paper notes verifiers sometimes fail at this kind of “variable binding”: attaching each quantity to the right thing.

Why it matters

A score at every step also tells you where a solution went wrong, not only whether it did. That is the seed of the process reward model, whose step labels come from people rather than from the final answer; the reasoning lesson's first_bad_step is the simplest version.

What changed since 2021

In the paperTodayLearn it
GSM8K is hard for the largest modelsClose to saturated; harder sets (MATH and beyond) took over as the frontier testbenchmarks
A verifier trained on final-answer labelsThe outcome reward model; step-level labels gave the process reward modelLet's Verify Step by Step companion
Best-of-100 at test timeOne of several ways to spend test-time compute, alongside votes, search and long chainspass@n
Lures fool the verifier past 400 samplesKnown as reward model over-optimisation; programs that check answers are preferred where they existverify
The generator never learns from the verifierThe check becomes the reward for reinforcement learning, which trains the generator itselfDeepSeekMath companion
Calculator annotations in the training textStructured tool calls to calculators and code interpreterstools

Glossary

Every term with hover guidance on this page, in one place.