Let's Verify Step by Step, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
- The step scorer lets you set how sure a reward model is about each step of a solution and watch the two ways of turning step scores into one solution score disagree.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. This paper is the sequel to Training Verifiers to Solve Math Word Problems (same last author, same idea of ranking sampled solutions); read that first if “verifier” or “best-of-N” is new. The reasoning models lesson builds both kinds of checker from scratch.
Abstract
“Our process-supervised model solves 78% of problems from a representative subset of the MATH test set.”Lightman et al. (2023), Abstract. Read the original
Everyday picture
Two teachers train an assistant marker. One only ever tells the assistant “this paper's final answer is right” or “wrong”. The other goes through papers line by line and says “fine, fine, fine, this line is where it goes wrong”. Which assistant becomes the better judge of new work? This paper runs that experiment at scale, on hard competition maths.
What the paper claims
- Process supervision wins. A process reward model (PRM), trained on human labels for each step, picks a correct solution out of 1,860 samples 78.2% of the time on its MATH test problems; an outcome reward model (ORM) manages 72.4%, and a majority vote 69.6%.
- A large reward model can stand in for human labellers when training small ones, which makes careful comparisons affordable.
- Active learning (showing labellers the solutions most likely to fool the current model) makes each label about 2.6 times as valuable.
- The data is public: PRM800K, 800,000 step-level human labels.
Why it matters today
PRM800K became a standard resource for training step checkers, and “score every step” became one of the main ideas for making reasoning reliable. The paper also sharpened a question that still runs through reasoning research: is it enough to reward right answers, or must the reasoning itself be checked?
1 Introduction · original
“It provides more precise feedback, since it specifies the exact location of any errors that occur.”Lightman et al. (2023), §1, on process supervision
Everyday picture
A long calculation is a chain: one broken link and everything after it is wrong, however careful the rest. Models that write long chains of thought sometimes invent a fact or make a slip in a moment of uncertainty (a hallucination), and in multi-step reasoning one slip is enough.
Tiny example: two ways to give feedback
The lesson's toy chain for 3 + 5 + 8 + 2: 3+5=8, 8+8=17, 17+2=19. Outcome supervision says one thing: the final answer 19 is wrong. Process supervision says: step 1 fine, step 2 wrong. Now take 3+5=9, 9+8=16, 16+2=18: two slips cancel, so outcome supervision calls it right, and process supervision still flags step 1. Checking steps catches right answers reached by wrong reasoning, which models trained on outcomes are known to produce.
Where reward models are used
A reward model that tells good outputs from bad can be used two ways: as the reward in a reinforcement-learning pipeline, or to search by rejection sampling, keeping the best of many samples. Either way, the system is only as reliable as the reward model, so this paper studies how to train the most reliable one.
Why this paper, when the question had been asked
Uesato et al. (2022) had already compared the two on grade-school maths (GSM8K) and found similar final performance. This paper changes three things: a more capable base model (GPT-4), much more human feedback, and the harder MATH dataset. Section 7 explains why the answers differ.
Why it matters
The introduction also gives an alignment argument: process supervision rewards a chain of thought people can follow and endorse, rather than any route that happens to end at the right answer. §6.2 returns to it.
2 Methods · original
“Outcome supervision can be provided without humans, since all problems in the MATH dataset have automatically checkable answers.”Lightman et al. (2023), §2
Everyday picture
Grading a final answer against a key is clerical work a program can do. Deciding whether line 7 of a proof is valid needs a person who knows the maths. That asymmetry shapes the whole paper: outcome labels are free, process labels are expensive.
Two regimes
- Large scale. Everything fine-tuned from GPT-4, with human process labels. Aim: the best ORM and PRM possible. Their training sets differ (§3 explains why), so this is not a controlled comparison.
- Small scale. Models trained with about 200 times less compute, where the large PRM plays the part of the human labeller. Aim: a controlled comparison and ablations that human labelling would make unaffordable (§4).
Why it matters
Separating “best possible result” from “fair comparison” is good experimental practice: the first shows what is achievable, the second shows why.
2.1–2.3 Scope, base models and the generator · original
Everyday picture
To compare two markers fairly, give them the same stack of scripts, written by the same student, and judge each marker only by how often the script it rates highest is actually right.
Tiny example: how a reward model is scored
For every test problem, sample N solutions from one fixed generator, let the reward model pick its favourite, and check that one's final answer. With 3 problems where the favourite is right, right and wrong, the reward model scores 2/3 = 67%. This is best-of-N, exactly as in Cobbe et al. (2021).
The setup
- Only the reward model is trained. The generator is never improved with reinforcement learning here; the paper calls that a natural next step and deliberately leaves it out.
- Base models. GPT-4 before any RLHF, and smaller models of similar design. All are first trained on MathMix, about 1.5B tokens of maths text, which improves mathematical reasoning.
- The generator writes one step per line. It is fine-tuned for one epoch on its own few-shot solutions that reached correct answers, only to teach the format: steps separated by newlines, so that a step is easy to point at.
Why it matters
Because the reward model is judged only by what it selects, the comparison measures exactly the thing you need from a checker: telling right from wrong among solutions that look alike. The verifier_experiment in the reasoning lesson uses the same protocol.
2.4 Data collection · original
“We expect to gain more information from labeling convincing wrong-answer solutions, since we know the PRM is mistaken about at least one step in each such solution.”Lightman et al. (2023), §2.4
Everyday picture
A tutor's time is precious. Showing them a solution with an obvious blunder in line 1 teaches your marker nothing it doesn't already know. Showing them a solution your marker loves but whose final answer is wrong is gold: somewhere in it is a mistake the marker cannot yet see.
Tiny example: the labelling task
The paper's Figure 1 shows the interface on a short algebra problem: the denominator of a fraction is 7 less than 3 times the numerator, and the fraction equals 2/5; what is the numerator? (The answer is 14.) Each line of the model's solution gets a label.
Hover or tap a step to see why it gets its label.
Reading it: the top box is the problem, with its known final answer, which labellers are shown. Below it, one row per step, in the order the model wrote them. The first five are correct and push towards the answer, so they are positive (green). Step 6 solves 5x = 6x − 14 wrongly: subtracting 5x from both sides gives x = 14, not 7. It is negative (pink). A labeller who finds that one bad step has told the reward model more than “wrong answer” ever could: where the solution broke.
Three labels, and what they mean
- Positive: correct and reasonable. Appendix D sharpens it: a step that would be neutral and makes progress towards the solution.
- Negative: incorrect or unreasonable.
- Neutral: §2.4 describes it as ambiguity (a subtly misleading step, or a poor suggestion that is technically valid); Appendix D's instructions define it as a step that is appropriate, correct and easy to verify but makes no progress. Either way, the choice of treating it as good or bad is postponed to test time.
Choosing what to label: convincing wrong answers
The labellers see solutions from the large generator only. Rather than labelling a random sample, the team surfaces convincing wrong-answer solutions: ones the current best PRM rates highly (convincing) whose final answer is wrong. “Wrong-answer” rather than “wrong”, because correctness is judged only by the final answer, which occasionally misgrades. Between batches the PRM is retrained on the latest labels, so what counts as convincing keeps moving. This is active learning; §4.2 measures how much it helps.
PRM800K
The result: about 800K step-level labels over 75K solutions to 12K problems. To have enough problems, 4.5K of the 5K MATH test problems went into training, leaving 500 test problems, chosen at random and checked to match the full test set's spread of subjects and difficulty (Appendix C). Every MATH result in the paper is on those 500.
Why it matters
Labelling cost is the bottleneck for any human-feedback method. Spending it where the model is confidently wrong is a pattern that recurs across machine learning, from spam filters to preference data.
2.5 Outcome-supervised reward models (ORMs) · original
Everyday picture
The marker from Cobbe et al. (2021), unchanged: trained on piles of solutions labelled right or wrong by their final answer, asked for one verdict at the end.
Tiny example
An ORM reads a 4-token solution and outputs, after each token, its running estimate that the solution is correct: 0.40, 0.55, 0.70, 0.62 (illustrative). Only the last one counts: the solution's score is 0.62.
The math: the ORM's score
In words: “an outcome reward model's score for a solution is its prediction at the solution's final token.”
With the numbers: the predictions are 0.40, 0.55, 0.70, 0.62, so T = 4 and the score is v4 = 0.62.
In Python:
# the ORM's prediction after each token (illustrative)
v = [0.40, 0.55, 0.70, 0.62]
T = len(v)
# score_ORM = v_T: the last token's prediction; Python counts from 0
v[T - 1] # → 0.62
It is trained exactly as Cobbe et al.'s token-level verifier: the same target (right or wrong) at every token, one epoch, with no dropout and no language-modelling loss this time (Appendix E). Training samples are drawn uniformly from the generator at temperature 1.0.
Why it matters
Outcome labels are cheap and plentiful, but noisy in one specific way: a solution that reaches the right answer through wrong reasoning is labelled correct, a false positive. The reasoning lesson's outcome_reward is the labelling rule.
2.6 Process-supervised reward models (PRMs) · original
“We define the PRM score for a solution to be the probability that every step is correct under the PRM.”Lightman et al. (2023), §2.6
Everyday picture
A bridge is only as safe as all of its joints together. If an inspector gives each joint a chance of being sound, the chance the whole bridge is sound is those chances multiplied: one doubtful joint drags the whole bridge down, however good the rest.
Tiny example
A PRM reads a 4-step solution and gives each step a chance of being correct: 0.98, 0.95, 0.97 and 0.10 (illustrative). The product is 0.98 × 0.95 × 0.97 × 0.10 ≈ 0.09. A 4-step solution with 0.96 on its last step instead scores 0.867. One bad step is enough to sink a solution.
Hover or tap a step. Start with the pink one in the middle: it is where the solution breaks.
Reading it: read down the solution. The first three steps are sound: substituting y = x² and factoring gives (x⁴ + 4)(x⁴ − 1). The fourth step claims x⁴ + 4 = (x² + 2)(x² − 2), but that product is x⁴ − 4, and the PRM scores it low (pink). Everything after inherits the error, and the solution also multiplies the factors' values where the problem asks for their sum, which the PRM flags too. The final answer 0 is wrong; the right factorisation, (x² + 2x + 2)(x² − 2x + 2)(x² + 1)(x + 1)(x − 1), gives 5 + 1 + 2 + 2 + 0 = 10. The paper's point: the PRM found the exact line.
The math: from step scores to a solution score
In words: “a solution's score is the chance that every step is correct: multiply together the PRM's probability for each step.”
With the numbers (illustrative): K = 4 and p = 0.98, 0.95, 0.97, 0.10: the product is 0.090. The alternative reduction, the minimum, gives 0.10: same verdict here. The two part ways on long solutions: 10 steps at 0.97 each multiply to 0.737 and 20 steps to 0.544, while the minimum stays 0.97. The product quietly favours shorter solutions, which the paper notes (Appendix F).
In Python:
import math
# p_k: the PRM's chance that step k is correct (illustrative)
p = [0.98, 0.95, 0.97, 0.10]
# Π_k p_k
round(math.prod(p), 3) # → 0.09
# the other reduction the paper tried
min(p) # → 0.1
# a solution whose last step looks fine instead
round(math.prod([0.98, 0.95, 0.97, 0.96]), 3) # → 0.867
# the product's length bias: 10 or 20 steps, each 0.97
round(0.97 ** 10, 3), round(0.97 ** 20, 3) # → (0.737, 0.544)
The math: what counts as a “correct” step
The PRM predicts one of three labels for each step. To get pk the paper adds the chance of “neutral” to the chance of “positive” (neutral counts as correct), or leaves it out.
In words: “a step's score is the PRM's probability that the step is positive, plus its probability that the step is neutral.”
With the numbers (illustrative): a step with P(+) = 0.70, P(neutral) = 0.20, P(−) = 0.10 scores 0.90 with neutral counted as positive, 0.70 without.
In Python:
# the PRM's three probabilities for one step (illustrative)
P_pos, P_neu, P_neg = 0.70, 0.20, 0.10
# neutral counts as correct (the paper's default)
round(P_pos + P_neu, 2) # → 0.9
# neutral counts as wrong
P_pos # → 0.7
| Neutral steps count as | Product | Minimum |
|---|---|---|
| positive | 78.2 | 77.6 |
| negative | 77.4 | 77.8 |
Reading it: each cell is one way of turning step predictions into a solution score. The spread is under a point: the product with neutral-as-positive is best (highlighted) and is used everywhere else in the paper, but how the step scores are combined matters much less than having them.
The math: supervising only up to the first mistake
When a solution is wrong, labellers stop at its first negative step. So a K-step solution whose first mistake is at step m yields labels for steps 1 to m only.
In words: “every step before the first mistake is labelled correct, the first mistake is labelled wrong, and nothing after it is labelled.”
With the numbers: in Figure 1's solution, K = 6 and the first mistake is m = 6, so the labels are 1, 1, 1, 1, 1, 0. Had the mistake been at step 3, the labels would be 1, 1, 0, and steps 4 to 6 would go unlabelled. A correct solution gets 1 on every step.
In Python:
def process_labels(K, m):
# m: the first wrong step, or None if every step is fine
if m is None:
return [1] * K
# y_k = 1 before m, 0 at m, and no label after m
return [1 if k < m else 0 for k in range(1, m + 1)]
process_labels(6, 6) # → [1, 1, 1, 1, 1, 0]
process_labels(6, 3) # → [1, 1, 0]
process_labels(6, None) # → [1, 1, 1, 1, 1, 1]
Why stop there? It keeps the comparison fair: on a correct solution both kinds of supervision say the same thing (all fine), and on a wrong one both say “there is a mistake”; process supervision adds only where. It also keeps labelling cost comparable, since without an answer key, checking a solution means finding its first mistake anyway.
Why it matters
The PRM is trained as an ordinary language model: after the last token of each step it predicts one token, the label, and training maximises the likelihood of the right one. No special architecture, and one forward pass scores every step. The reasoning lesson's first_bad_step and check_step are a PRM with a perfect step checker.
Try it: score a solution
Everyday picture
You are the PRM. For each of Figure 1's six steps, set how sure you are that it is correct, and watch the solution's score.
Tiny example
The sliders start at illustrative values: five steps you trust (0.95 to 0.99) and step 6, the wrong one, at 0.08. The product is about 0.07; the minimum 0.08. Now drag step 6 up to 0.9 and both climb: a PRM that misses the bad step would promote this wrong solution.
Try it: move the sliders. Watch how one low step controls the product, and how shaving every step to 0.9 lowers the product (0.53) while the minimum stays at 0.9.
Reading it: each row is one step of the solution in Figure 1, with a slider for the PRM's probability that it is correct. The bar under the sliders shows the running product, step by step, as a share of 1: it can only shrink as it moves down the solution. The readout gives the two solution scores the paper compared (product and minimum) and the first step you rated below 0.5, which is where a labeller would stop. Every number is yours; nothing here comes from a trained model.
Why it matters
The product and the minimum agree when one step is clearly bad, and disagree when many steps are each a little doubtful. Table 4 says the choice barely matters for this PRM; for a checker whose step scores drift down on long solutions, the length bias of the product can.
3 Large-scale supervision · original
“Not only does the PRM reach higher performance for all values of N, but the performance gap widens as N increases.”Lightman et al. (2023), §3
Everyday picture
Three ways to pick one solution from a pile: trust the answer most solutions agree on (a vote), trust the marker who only grades answers (ORM), or trust the marker who checks every line (PRM). The bigger the pile, the more the choosing rule matters.
Tiny example: why the training sets differ
The PRM trains on PRM800K. The ORM could train on PRM800K's solutions too, but active learning made them mostly wrong-answer solutions (86% of them end in a wrong answer; Appendix B), a poor diet for learning what right looks like. So the ORM gets 100 uniform samples per problem instead: about ten times more data and no overlap. Mixing uniform samples into PRM800K did not help the ORM.
Hover the chart, or tab to it and use the arrow keys, to read each method at each N.
Reading it: the x-axis is N, the number of generator solutions per problem, on a log scale from 10 to 1,860; the y-axis is the share of the 500 test problems where the chosen solution is right. At N = 10 the three methods are close (PRM 67.8%, ORM 67.0%, vote 63.3%). As N grows the vote levels off at about 69.6% and the ORM at about 72.4%, while the PRM keeps climbing to 78.2%. A better checker can use a bigger pile; a weaker one can't. Redrawn from the paper's Figure 3; values recovered from the plotted points, and the end points match the paper's printed 78.2, 72.4 and 69.6. (The paper shades the spread over many subsamples of the 1,860 solutions; the redraw shows the mean.)
| Method | % solved |
|---|---|
| Majority voting | 69.6 |
| Outcome-supervised reward model | 72.4 |
| Process-supervised reward model | 78.2 |
Why it matters
Majority voting (self-consistency) needs no training and is a strong baseline; here both reward models beat it, the PRM by 8.6 points. Voting weighted by reward-model scores, which combines the two, did not noticeably help. Compare Cobbe et al., whose verifier stopped improving at 400 samples: here the ORM is flat from a few hundred on, while the PRM is still climbing at 1,860. Appendix G shows the gap at every difficulty level, and on the easiest fifth of problems the ORM even gets slightly worse as N grows (it finds solutions that fool it) while the PRM does not.
4 Small-scale synthetic supervision · original
“This setup enables us to simulate a large amount of data collection at a modest cost.”Lightman et al. (2023), §4
Everyday picture
You want to know whether line-by-line feedback teaches better than answer-only feedback, but hiring tutors for every variant is unaffordable. So you let your best trained marker play tutor for a class of junior markers, and give some juniors its line-by-line verdicts and others only its overall verdicts.
Tiny example: turning the large PRM into a labeller
The large PRM (PRMlarge) gives each step a probability of being negative, for example 0.03, 0.12, 0.35, 0.60 (illustrative). Any step above 0.2 counts as a mistake, because PRMlarge leans slightly towards calling steps positive (Appendix H). Step 3 is the first above 0.2.
The math: synthetic labels
In words: “the first mistake is the earliest step that the large PRM thinks is negative with more than 20% probability.”
With the numbers (illustrative): 0.03 and 0.12 are below 0.2, 0.35 is above, so m = 3. Process supervision gives the small model labels 1, 1, 0 (up to the first mistake, as with people). Outcome supervision says only: not every step is correct, so the solution is wrong (0). A third kind, outcome supervision from final-answer checking, is also trained for comparison.
In Python:
# PRM_large's probability that each step is negative (illustrative)
P_neg = [0.03, 0.12, 0.35, 0.60]
threshold = 0.2
# m: the first step above the threshold, counting from 1
m = next(k for k, q in enumerate(P_neg, 1) if q > threshold)
m # → 3
# process supervision: labels up to and including the first mistake
[1 if k < m else 0 for k in range(1, m + 1)] # → [1, 1, 0]
# outcome supervision from PRM_large: correct only if no step is flagged
int(all(q <= threshold for q in P_neg)) # → 0
Why it matters
Using a strong model as a stand-in labeller (Gao et al., 2022 did the same for preference data) turns a question that would cost a fortune in human labour into a cheap, repeatable experiment, at the price of trusting the stand-in.
4.1 Process vs outcome supervision · original
Everyday picture
Same scripts, same junior markers, only the kind of feedback differs. If the line-by-line group ends up better, the feedback is the reason.
Tiny example
The striking pair of numbers in Figure 4(a): a small PRM trained on one labelled solution per problem selects correctly 48.6% of the time; a small ORM trained on final-answer checks needs 200 solutions per problem to reach 48.5%.
Reading it: in panel (a) the x-axis (log scale) is how many solutions per problem were labelled to train each small reward model, and the y-axis is its best-of-500 score. The PRM line sits above both ORM lines at every data size, and the gap does not close: 58.6% against 52.4% and 48.5% at 200 labelled solutions per problem. The active-learning line sits higher still; it comes in §4.2. In panel (b) the x-axis is N at test time, for the best model of each kind: all three start near 22% at N = 1, then the PRM pulls away, reaching 60.9% at N = 1,000 against 52.8% and 48.5%. Redrawn from the paper's Figure 4; values recovered from the plotted points (means over three seeds).
Which ORM is the fair baseline?
Final-answer checking mislabels right-answer, wrong-reasoning solutions, and MATH has many of those, so it may exaggerate the ORM's handicap. Outcome labels from PRMlarge avoid that (it judges the reasoning, then reports only a verdict), and the ORM trained on them does better in panel (b). The paper considers that the fairer baseline, while inviting readers to draw their own conclusions. Either way the PRM wins.
Why it matters
This is the controlled experiment §3 could not be: identical solutions, only the form of supervision changed. The answer is that knowing where the mistake is makes a reward model better at every data size tested.
4.2 Active learning · original
Everyday picture
If you can only afford to have ten scripts marked, pick the ten your marker is most wrongly confident about, plus a couple of its favourites so it still sees what good work looks like.
Tiny example: picking 5 from a pool of 8
A small selector PRM, trained on one sample per problem, scores the pool (illustrative): a 0.97 right answer, b 0.93 wrong, c 0.90 wrong, d 0.85 right, e 0.71 wrong, f 0.64 wrong, g 0.40 wrong, h 0.22 right. To label N = 5: 80% of them (4) are the most convincing wrong-answer solutions, b, c, e and f; the other 20% (1) is the most convincing of the rest, a.
The math: the 80/20 selection
In words: “of the N solutions labelled per problem, four in five are the selector's highest-rated wrong-answer solutions, and the rest are its highest-rated remaining solutions, right or wrong.”
With the numbers: N = 5 gives nwrong = 4 and nrest = 1, selecting b, c, e, f and then a.
In Python:
# (name, selector score, reached the right answer?), illustrative
pool = [("a", 0.97, True), ("b", 0.93, False), ("c", 0.90, False), ("d", 0.85, True),
("e", 0.71, False), ("f", 0.64, False), ("g", 0.40, False), ("h", 0.22, True)]
N = 5
n_wrong = round(0.8 * N)
ranked = sorted(pool, key=lambda s: -s[1])
wrong = [s for s in ranked if not s[2]][:n_wrong]
rest = [s for s in ranked if s not in wrong][:N - n_wrong]
n_wrong, [s[0] for s in wrong], [s[0] for s in rest] # → (4, ['b', 'c', 'e', 'f'], ['a'])
The chosen solutions are then labelled by PRMlarge, and a new small PRM trains on them. The mix guarantees every selected solution is convincing, most contain a known mistake, and the set is not entirely wrong answers.
The math: how much is it worth?
Look at panel (a) of Figure 4 again. On its log x-axis, the uniform PRM and the active-learning PRM are nearly parallel straight lines. Parallel lines on a log axis are a fixed ratio apart: one reaches any given score with a constant multiple of the other's data.
In words: “fit each line as a score = a + b × log₁₀(data); the horizontal gap between them, in decades, is the difference in intercepts divided by the slope, and ten to that power is how many times more data the uniform method needs.”
With the numbers: fitting the plotted points (without the active-learning point at 200, which the paper says falls below trend) gives aU = 48.67 and b = 4.39 for uniform labelling, aAL = 50.60 for active learning. The gap is 1.93 / 4.39 = 0.44 decades, and 100.44 ≈ 2.7: close to the paper's 2.6, which comes from its own fit to its exact data.
In Python:
import math
def fit(xs, ys):
# least squares for y = a + b · log10(x)
x = [math.log10(v) for v in xs]
x_bar, y_bar = sum(x) / len(x), sum(ys) / len(ys)
b = sum((xi - x_bar) * (yi - y_bar) for xi, yi in zip(x, ys)) / sum((xi - x_bar) ** 2 for xi in x)
return round(y_bar - b * x_bar, 2), round(b, 2)
# Figure 4(a), values recovered from the plot: solutions labelled per problem, % solved
a_U, b = fit([1, 2, 5, 10, 20, 50, 100, 200], [48.56, 49.9, 51.87, 53.02, 54.71, 56.14, 57.48, 58.59])
a_U, b # → (48.67, 4.39)
a_AL, b_AL = fit([2, 3, 6, 11, 21, 51, 101], [51.23, 53.46, 53.94, 55.14, 56.36, 58.17, 59.17])
a_AL, b_AL # → (50.6, 4.36)
# the horizontal gap on the log axis, as a data multiple
round(10 ** ((a_AL - a_U) / b), 1) # → 2.8
(Rounded fits give 2.8; unrounded, 2.74. Either way, close to the paper's 2.6.)
Two honest caveats from the paper
- The largest active-learning set (200 per problem, out of a pool of 1,000) falls a little below the trend, probably because 200 of 1,000 leaves little room to be selective.
- Retraining the selector during collection, as the real PRM800K collection did, was unstable in these small-scale tests and gave no gain. The team expects some form of it to help but has no evidence yet.
Why it matters
A 2.6× saving on the most expensive part of the pipeline is a lot. The idea, spend labels where the current model is confidently wrong, is how the human labels in PRM800K were chosen.
5 Out-of-distribution generalization · original
“This shows us that the PRM can tolerate a modest amount of distribution shift and that its strong performance holds up on fresh test questions.”Lightman et al. (2023), §5
Everyday picture
A marker trained on one exam board's papers is handed a different board's newest papers, written after the marker's training ended. Does its judgement still hold?
Tiny example: the test
Recent AP Calculus, AP Chemistry, AP Physics, AMC10 and AMC12 questions, released after the base model's pretraining data was collected, so it cannot have seen them: out-of-distribution and uncontaminated. Each method picks from 100 samples per problem.
| Exam | ORM | PRM | Majority vote | Problems |
|---|---|---|---|---|
| AMC10/12 | 49.1 | 53.2 | 32.8 | 84 |
| All exams | 63.8 | 72.9 | 61.3 | 234 |
Reading it: each row is an exam, each column a way of choosing among 100 samples. On every exam the PRM column is highest, as on MATH. The competition problems (AMC) are hardest, and there majority voting falls far behind both reward models: when most samples are wrong, the most common answer is often wrong too.
The math: an aggregate is a weighted average
In words: “the overall score is each exam's score weighted by how many problems it has, divided by the total number of problems.”
With the numbers: the four rows of Table 1 have 45, 60, 45 and 84 problems: 234 in all. Rebuilding the aggregates from the rows gives 73.0 for the PRM and 61.4 for majority voting, within rounding of the printed 72.9 and 61.3. For the ORM it gives 63.5, not the printed 63.8, a gap too large for rounding.
In Python:
# Table 1 rows: problems n_e, then % for ORM, PRM, majority vote
rows = {"AP Calculus": (45, 68.9, 86.7, 80.0), "AP Chemistry": (60, 68.9, 80.0, 71.7),
"AP Physics": (45, 77.8, 86.7, 82.2), "AMC10/12": (84, 49.1, 53.2, 32.8)}
total = sum(n for n, *_ in rows.values())
total # → 234
def aggregate(col):
# Σ_e n_e · acc_e / Σ_e n_e
return round(sum(r[0] * r[col] for r in rows.values()) / total, 1)
aggregate(1), aggregate(2), aggregate(3) # → (63.5, 73.0, 61.4)
Two small inconsistencies, then: the text says the held-out set has 224 questions, while the table's rows add up to 234; and the ORM's aggregate does not quite follow from its rows. Neither changes the conclusion.
Why it matters
A reward model that only works on its training distribution is of limited use. This is a small test (234 problems), but it points the same way as the main result on problems the model cannot have memorised.
6 Discussion · original
Everyday picture
Three questions a sceptical reader would ask: why does step feedback help, is it good for anything besides accuracy, and could the model simply have seen the test?
The three answers
Credit assignment (§6.1), alignment (§6.2), and contamination (§6.3), each below.
Why it matters
The first gives a mechanism, the second a motive beyond benchmarks, the third a check that the numbers mean what they seem to.
6.1 Credit assignment · original
“Process supervision makes credit assignment easier, and we believe that this explains its strong performance.”Lightman et al. (2023), §6.1
Everyday picture
A team loses a match. “You lost” is feedback, but it doesn't say who fumbled. On hard problems nearly every attempt is wrong somewhere, so “you lost” arrives almost every time and says almost nothing. That is the credit assignment problem.
Tiny example: how much a process label adds
A wrong 6-step solution. Outcome supervision says “wrong”. Process supervision also says which of the 6 steps is the first mistake. Picking one of 6 equally likely positions is log2 6 ≈ 2.6 bits of extra information; for a 20-step solution, 4.3 bits. (This counting is this page's illustration of the paper's argument, not a calculation from the paper.)
The math: the extra information
In words: “telling someone which one of K equally likely places the first mistake is in gives them log base 2 of K bits.”
With the numbers: K = 6 gives 2.58 bits; K = 20 gives 4.32. Longer solutions, where outcome labels are least informative, gain the most.
In Python:
import math
# which of K steps holds the first mistake: log2 K bits
round(math.log2(6), 2) # → 2.58
round(math.log2(20), 2) # → 4.32
Why it matters
An ORM has to work out on its own where wrong solutions went wrong, from nothing but final verdicts; the paper argues this is hardest exactly on hard problems. A PRM is told. The reasoning lesson's diagram of an outcome verifier against a process verifier (§4 of the lesson) is this contrast on a toy chain.
6.2 Alignment impact · original
“Our results show that process supervision in fact incurs a negative alignment tax.”Lightman et al. (2023), §6.2
Everyday picture
Usually, making a system safer costs some performance, like a speed limiter on a car. Occasionally the safer design is also the faster one. Then there is no argument left against it.
Tiny example
An alignment tax is the accuracy you give up for a safer training method. Here the safer method (reward reasoning people can check) scores 78.2% against 72.4%: the tax is 72.4 − 78.2 = −5.8 points, negative, a bonus.
The argument
- Interpretability. Process supervision rewards a chain of thought that follows a process people endorse, so the reasoning is more likely to be readable.
- Safety. It rewards the aligned reasoning directly, instead of using outcomes as a proxy. Outcome rewards can, in the worst case, teach a model to exploit the reward signal (reward hacking).
- Adoption. Safer methods that cost performance are hard to adopt under pressure to deploy the most capable model; one that helps performance has no such obstacle. The paper notes the result is for maths, and other domains need testing.
Why it matters
The argument concerns reward models used to train a policy, which this paper never does. It is a motivation, not a measured result: the measured part is the negative tax on best-of-N selection.
6.3 Test set contamination · original
Everyday picture
If a student has seen the exam paper, their score says little. MATH problems are discussed online, so some may be in the pretraining data.
Tiny example: the evidence the paper offers
- MathMix was filtered for MATH problems by string matching, but rephrased copies can slip through, so no strong guarantee is possible.
- Inspecting solutions showed no clear signs of memorisation.
- The PRM often finds correct solutions to problems the generator solves only a few percent of the time (Appendix I shows pass rates as low as 0.1%): a model that had memorised a problem would solve it often.
- The out-of-distribution exams of §5, which are guaranteed unseen, show the same pattern.
Why it matters
Contamination would lift every method alike, so the comparisons should survive even if absolute numbers are slightly inflated. Benchmark contamination is covered in the benchmarks lesson.
7 Related work · original
“The trend also shows that process supervision beats outcome supervision when scaled up, even when judged based solely on outcomes.”Lightman et al. (2023), §7.1
Everyday picture
Two studies seem to disagree: one found line-by-line and answer-only feedback about equally good, this one finds line-by-line clearly better. Both can be right if they measured at different amounts of feedback.
Tiny example: reconciling Uesato et al.
In Figure 4(a), a PRM trained on 1 labelled solution per problem (48.6%) matches an ORM trained on 200 (48.5%). So a little process supervision and a lot of outcome supervision can tie, which is roughly the regime Uesato et al. (2022) measured. Scale the process supervision up and it pulls ahead. Their other findings fit too: process supervision reached the same performance with less data.
The other threads
- Synthetic supervision. Gao et al. (2022) used a large “gold” reward model in place of people to study over-optimisation in RLHF; §4 borrows the move.
- Reasoning in language. Training on technical text (Minerva), self-consistency, chain of thought and scratchpads, and zero-shot “let's think step by step”, which the paper's title echoes.
Why it matters
A result that appears to contradict earlier work is most convincing when it explains the earlier result instead of dismissing it. Here one data-scaling curve does both.
8 Conclusion · original
Everyday picture
Line-by-line marking produces better markers, and you can make it affordable by marking only the scripts your marker gets confidently wrong.
The claims, in one place
- Process supervision trains much more reliable reward models than outcome supervision for mathematical reasoning (78.2% against 72.4% best-of-1860).
- Active learning lowers the cost of human labels, by about 2.6× in the small-scale test.
- PRM800K is released, with the hope that removing the cost of collecting such data speeds up research on aligning language models.
Why it matters
The authors call process supervision “under-explored” and look forward to seeing how far it generalises beyond maths. The open questions they leave (does a PRM make a better training reward than an ORM? does it work where answers can't be checked?) are the ones the field took up next.
Appendix B: inside PRM800K · original
Everyday picture
A dataset built by active learning is deliberately lopsided: it is mostly the hard cases, because those are the ones worth paying for.
Tiny example: two phases
- Phase 1 (about 5% of the data, around 40K labels): labellers rated several alternative next steps at each point and could write their own; the solution continued from a positive one. Slow, and prone to long runs of neutral steps.
- Phase 2 (the rest, in 10 generations): whole solutions were pre-generated, the current best PRM picked the most convincing wrong-answer ones, labelling stopped at the first negative step, and the PRM was retrained between generations.
| Phase 1 | Phase 2 | Combined | |
|---|---|---|---|
| % of solutions ending in a correct answer | 85.1 | 13.2 | 14.2 |
| % of steps labelled correct | 58.6 | 74.1 | 73.1 |
Reading it: read the columns left to right. Phase 1 solutions mostly end correctly (85.1%), because labellers steered them. Phase 2 was chosen to be wrong, and only 13.2% end correctly. Yet within those wrong solutions most steps are fine (74.1%): a wrong solution is usually right until its first mistake, and labelling stops there. So even a dataset of mostly wrong solutions teaches a lot about what correct steps look like.
Quality control
Before phase 2, labellers had to agree with the researchers' gold labels on at least 75% of 30 screening questions; 10 to 20 more quality-control problems per generation were mixed into their work, and labellers whose quality slipped were removed. In all, 1,085,590 step labels were collected on 101,599 solutions; dropping quality-control and incomplete labels leaves the roughly 800K used for training. The full set is released at github.com/openai/prm800k.
Why it matters
Human labels are never perfect: the DeepSeekMath companion notes that later work found a noticeable share of PRM800K's labels to be wrong. Screening, spot checks and agreement rates, as in the fine-tuning lesson's label_agreement, are how any labelling project keeps that share down.
What changed since 2023
| In the paper | Today | Learn it |
|---|---|---|
| PRM used only to rank samples (best-of-N) | Step-level scores also used to guide search and as training rewards, with mixed results as rewards: learned step judges can be gamed | reasoning |
| 800K human step labels | Step labels generated automatically, for example by estimating how often continuations from each step reach the right answer | Training Verifiers companion |
| Outcome supervision labelled the weaker method | Outcome rewards from programs that check answers became the workhorse of reasoning training, since a rule cannot be fooled the way a learned model can | DeepSeekMath companion |
| Convincing wrong answers surfaced for labelling | A standard pattern for spending labelling budgets where the model is confidently wrong | fine-tuning |
| Majority voting as the baseline | Still the baseline every checker must beat | self-consistency companion |
Glossary
Every term with hover guidance on this page, in one place.