An annotated companion · AI Primer

DeepSeekMath and GRPO, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The equations are reproduced with every symbol decoded; a few selected rows of the paper's tables appear with attribution. The paper's figures are redrawn from scratch, and any numbers made up for a worked example are labelled illustrative. Every section links to the original: read it alongside.

How to read this page

Nothing on this page assumes you already know the jargon.

  • Any dotted word explains itself when you hover it, tab to it, or tap it. So does every symbol in every equation.
  • Each equation is followed by a table of its symbols, a sentence reading it aloud, the equation worked on real numbers, and the same numbers computed in plain Python.
  • The diagrams are live: tap a block for its note, drag a slider and watch the numbers move.

The paper has two halves. The first builds a strong maths model from web data. The second, and the reason this paper is famous, introduces GRPO, the reinforcement learning method later used to train reasoning models. If you are here for GRPO, jump to section 4; it assumes only the PPO companion or the reinforcement learning lesson.

The running example

Section 4 uses one tiny case throughout. The question is “What is 12 × 3?”. The model writes a group of G = 4 answers, and a reward model scores them 0.9, 0.1, 0.3 and 0.7. Answers 1 and 4 say 36 (right); answers 2 and 3 say 15 and 38 (wrong). These scores are illustrative, chosen so the arithmetic stays checkable by hand.

Abstract · original

“Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.”Shao et al. (2024), Abstract

Everyday picture

To make a small model good at maths, the authors did two things. They gathered a mountain of maths from the public web, like a librarian who learns to spot maths books by their covers and then sweeps a city's worth of second-hand shops. Then they coached the model with practice problems, marking each attempt not against an absolute standard but against the model's other attempts at the same problem.

What the paper claims

  • DeepSeekMath 7B continues pretraining a 7-billion-parameter code model on 120B maths tokens from Common Crawl, plus natural language and code.
  • It scores 51.7% on the competition-level MATH benchmark with no tools and no voting, approaching Gemini-Ultra and GPT-4; self-consistency over 64 samples reaches 60.9%.
  • Two factors explain it: a carefully engineered pipeline for selecting web data, and GRPO, a variant of PPO that needs less memory.

Why it matters today

GRPO outgrew this paper: DeepSeek-R1 (2025) used it, with rule-based rewards, to train a model that writes long, self-checking chains of thought. The reasoning lesson traces how comparing answers within a group, with a checkable answer as the only reward, trains a model to think longer.

1 Introduction · original

“GRPO foregoes the critic model, instead estimating the baseline from group scores, significantly reducing training resources.”Shao et al. (2024), §1

Everyday picture

In early 2024 the best maths models (GPT-4, Gemini-Ultra) were closed, and open models trailed far behind. The paper's bet is that an open 7B model can close much of that gap if it is fed the right data and then trained with a cheap, well-aimed form of reinforcement learning.

The pipeline, and what each stage adds

StageModelGSM8KMATH
Continued pretraining on the new corpus (§2)DeepSeekMath-Base 7B64.2%36.2%
Instruction tuning (§3)DeepSeekMath-Instruct 7B82.9%46.8%
GRPO (§4)DeepSeekMath-RL 7B88.2%51.7%

All numbers are from the paper's §1 and Tables 2 and 5, with the model writing a chain of thought and no tools. GSM8K is grade-school word problems; MATH is competition problems. The paper also contributes (§1.1) evidence that code training before maths training helps, that arXiv papers surprisingly do not, and a “unified paradigm” that puts SFT, rejection sampling, DPO, PPO and GRPO in one formula (§5.2). It evaluates (§1.2) on English and Chinese maths benchmarks from grade school to college, on formal theorem proving, and on general language, reasoning and coding benchmarks.

Why it matters

The base model beating the much larger Minerva 540B (§2.3) is the paper's argument that data quality can stand in for parameters; the RL stage is its argument that a cheaper RL algorithm is enough to add several more points.

2 Math pre-training · original

2.1 Data collection and decontamination · original

“After four iterations of data collection, we end up with 35.5M mathematical web pages, totaling 120B tokens.”Shao et al. (2024), §2.1

Everyday picture

You want every maths book in a city of second-hand shops, but you only own a small shelf of known maths books. So you learn what maths books look like from your shelf, sweep the shops for look-alikes, and notice which shops turned out to be full of maths. Then you walk those shops' maths aisles by hand, adding what your eye missed to your shelf, relearn, and sweep again.

Tiny example

The domain rule works like this. Suppose a website has 1,000 pages and the first sweep collected 150 of them: 15% is over the paper's 10% threshold, so the whole domain is marked maths-related, and people then mark which URL paths hold maths (the paper's example: mathoverflow.net/questions). Pages under those paths that the classifier missed join the seed set for the next round. (The page counts here are illustrative; the 10% threshold is the paper's.)

repeat: 4 iterations Math seedOpenWebMath at first 1. Train a fastText model500K seed vs 500K random pages CC40B 2. Recall math-related pageskeep the top-scoring pages 3. Discover math-related domainsover 10% of pages collected 4. Annotate math URL pathsmissed pages join the seed

Hover or tap a step. Start at the top with the seed.

Figure 2 of the paper, redrawn: the iterative pipeline that collects mathematical web pages from Common Crawl. Based on Shao et al. (2024), Figure 2.

Reading it: read top to bottom, then follow the wire on the left back to the top. The seed shelf trains a small, fast text classifier; the classifier scores the whole deduplicated crawl (the box on the right, 40B HTML pages) and the best-scoring pages are kept. Step 3 turns page-level hits into site-level knowledge, and step 4 is the only place people are involved: they mark which parts of a maths-heavy site really hold maths, and those pages widen the seed so the next classifier recognises more kinds of maths. The loop ran four times; by the fourth round nearly 98% of what it found had already been found in the third, so collection stopped.

Decontamination

A corpus scraped from the web will contain benchmark questions, and training on them would inflate the scores (benchmark contamination). The paper removes any text segment containing a 10-word run that exactly matches any benchmark text (GSM8K, MATH, CMATH, AGIEval among them); for benchmark texts shorter than 10 words but at least 3, it removes pages containing an exact match. The ngram_overlap function in the benchmarks lesson builds this kind of n-gram overlap check.

Why it matters today

This is the quality classifier pattern from the pretraining lesson (its QualityClassifier is a naive Bayes version of the same idea), turned into a loop. The paper notes the recipe is not specific to maths: the same loop could gather code or any other domain.

2.2 Validating the quality of the corpus · original

Everyday picture

To judge a new cookbook, cook the same meal from each competing cookbook with the same cook and the same oven. The paper trains the same small model (DeepSeek-LLM 1.3B) for 150B tokens on each maths corpus and compares the results.

What the paper reports

Selected columns of Table 1, reproduced with attribution (Shao et al., 2024). DeepSeek-LLM 1.3B after 150B tokens on each corpus, few-shot chain-of-thought
Corpus (size in tokens)GSM8KMATHCMATH
No math training2.9%3.0%12.3%
MathPile (8.9B)2.7%3.3%1.2%
OpenWebMath (13.6B)11.5%8.9%16.8%
Proof-Pile-2 (51.9B)14.3%11.2%19.9%
DeepSeekMath Corpus (120.2B)23.8%13.6%41.5%

Reading it: each pair of bars is one corpus: the solid bar is GSM8K accuracy and the striped bar is MATH accuracy, both on a 0 to 30% scale. The DeepSeekMath Corpus leads on both. MathPile, which is mostly arXiv papers, barely moves the model from where it started, an early sign of the arXiv result in §5.1.2. The Chinese column in the table (CMATH, 41.5% against at most 19.9%) shows the second advantage the paper claims: the corpus is multilingual, while the others are mostly English and can even hurt Chinese maths. The paper's Figure 3 adds the third: the other corpora are small enough to be repeated several times in 150B tokens, and their learning curves flatten early.

Why it matters

Comparing corpora by training the same small model on each is how data teams decide what to keep; the pretraining lesson covers why the mixture and the number of repeats matter.

2.3 Training and evaluating DeepSeekMath-Base 7B · original

Everyday picture

Take a programmer who already reads code fluently and send them on a long maths course that still includes some programming and some ordinary reading, so they do not forget what they knew.

The recipe

Start from DeepSeek-Coder-Base-v1.5 7B and train for 500B tokens with this mixture:

The 500B-token training mixture of DeepSeekMath-Base 7B, drawn from the percentages in Shao et al. (2024), §2.3. Token counts are those percentages of 500B.

Reading it: each bar is one source's share of the 500B-token budget, scaled so the largest fills the row. More than half is the new maths corpus; a fifth is GitHub code, which keeps the starting model's coding skill alive; arXiv papers and natural-language web text (English and Chinese) take a tenth each; the shortest bar is AlgebraicStack, maths-flavoured code. The training settings follow the 1.3B experiments, with a peak learning rate of 4.2 × 10−4 and a batch of 10M tokens.

What the paper reports

Selected rows and columns of Tables 2 and 3, reproduced with attribution (Shao et al., 2024)
ModelGSM8KMATHMATH, Python
Minerva 540B (closed)58.8%33.6%n/a
Mistral 7B40.3%14.3%18.2%
Llemma 34B54.0%25.3%26.3%
DeepSeekMath-Base 7B64.2%36.2%31.4%

Three kinds of test: solving with a written chain of thought; solving by writing a Python program whose result is the answer (program-of-thought); and informal-to-formal proving on miniF2F in Isabelle, where it scores 24.6% on the test set against 21.3% for Llemma 34B (Table 3). On MATH the 7B base model beats Minerva 540B, a closed model 77 times larger. It also improves on its starting checkpoint on general reasoning (MMLU 54.9% and BBH 59.5%, against 42.9% on each for the checkpoint it was trained from, Table 4), while keeping most of its coding ability.

Why it matters

A 7B model is cheap to sample from many times, which is exactly what the RL stage of section 4 needs: 64 answers per question.

3 Supervised fine-tuning · original

Everyday picture

After the maths course comes a workbook of solved examples in the house style: show your working, or write a program, or mix the two. Supervised fine-tuning trains the model to imitate those worked solutions.

The data and the training (§3.1, §3.2)

  • 776K training examples in English and Chinese, from grade school to college, each solution written as a chain of thought, as a program (program-of-thought), or as tool-integrated reasoning. English sources include GSM8K and MATH problems annotated with tool-integrated solutions, a subset of MathInstruct and Lila-OOD; the Chinese set covers K-12 problems across 76 sub-topics.
  • Examples are packed together up to 4K tokens; 500 steps at batch size 256 and a constant learning rate of 5 × 10−5.
  • Result, DeepSeekMath-Instruct 7B: 46.8% on MATH with a chain of thought (above every open model the paper compares, including 70B ones) and 57.4% when it may run Python (Table 5).

Why it matters

This model is the starting point for GRPO, and the reference model the RL stage stays tethered to. The training stages lesson builds the SFT loss, sft_loss, from scratch.

4 Reinforcement learning · original

This is the part of the paper that travelled. Section 4.1.1 starts from PPO, shows what it costs, and removes the expensive piece.

4.1.1 From PPO to GRPO · original

“As the value function employed in PPO is typically another model of comparable size as the policy model, it brings a substantial memory and computational burden.”Shao et al. (2024), §4.1.1

Everyday picture

PPO grades each answer against a prediction of how well the model should do on that question, and the prediction comes from a second model, the value network, as big as the first. It is like hiring a second examiner whose only job is to guess each question's difficulty. GRPO fires the second examiner. It asks the same question several times and grades each answer against the others: on a question where most attempts succeed, success earns little; where most fail, the one success stands out.

PPO, written for a language model

Tiny example. An answer o has 2 tokens. Since the answer was sampled, training has made the first token 1.1 times as likely and the second 1.4 times as likely, and the answer's advantage is 0.5 at both tokens. With ε = 0.2 the second ratio counts only up to 1.2, so the two token terms are 1.1 × 0.5 = 0.55 and min(1.4 × 0.5, 1.2 × 0.5) = 0.6, and the answer's objective is their average, 0.575.

In words: “draw a question, let the old policy write an answer, and for each token of the answer take PPO's clipped, pessimistic ratio-weighted advantage; average over the tokens, and over many questions.”

With the numbers: (min(1.1 × 0.5, 1.1 × 0.5) + min(1.4 × 0.5, 1.2 × 0.5)) / 2 = (0.55 + 0.6) / 2 = 0.575.

In Python:

def clip(x, lo, hi):
    return max(lo, min(x, hi))
eps = 0.2
# π_θ(o_t | q, o_<t) / π_θold(o_t | q, o_<t) for the answer's two tokens, and each token's advantage
ratios, A = [1.1, 1.4], [0.5, 0.5]
terms = [min(r * a, clip(r, 1 - eps, 1 + eps) * a) for r, a in zip(ratios, A)]
[round(x, 3) for x in terms]  # → [0.55, 0.6]
# (1/|o|) Σ_t
round(sum(terms) / len(terms), 3)  # → 0.575

This is the clipped objective of the PPO companion, with one change of vocabulary: the state is the question plus the tokens written so far, and each action is one token. The advantage At comes, as in PPO, from generalized advantage estimation over the rewards and a learned value function Vψ.

The per-token KL penalty in PPO's reward

Tiny example. The reward model gives a finished answer 1.0, and it scores only the last token. At that last token the policy gives probability 0.5 where the frozen reference gave 0.4. With β = 0.04, the reward used for training at that token is 1.0 − 0.04 × ln(0.5 / 0.4) = 1.0 − 0.04 × 0.223 = 0.991. Every other token gets only its own KL fee.

In words: “the reward at each token is the reward model's score, minus a fee proportional to how much more likely the policy has made that token than the reference model did.”

With the numbers: 1.0 − 0.04 × ln(1.25) = 1.0 − 0.0089 = 0.991.

In Python:

import math
r_phi, beta = 1.0, 0.04
pi_theta, pi_ref = 0.5, 0.4
# r_t = r_φ − β log(π_θ / π_ref)
round(r_phi - beta * math.log(pi_theta / pi_ref), 3)  # → 0.991

The paper names two problems with this setup. The value function is typically another model as large as the policy, a heavy memory and compute burden. And since the reward model usually scores only the last token, it is hard to train a value function that is accurate at every token.

Figure 4, redrawn: PPO against GRPO

PPO KL score r v q Policy modeltrained o Referencefrozen Rewardfrozen Valuetrained + GAE A GRPO q Policy modeltrained Referencefrozen: KL in the loss o1o2…oG Reward modelfrozen r1r2…rG Group: (r − mean) / std A1A2…AG solid boxes are trained; dashed boxes are frozen

Hover or tap a block. Compare the two halves: find the box GRPO no longer has.

Figure 4 of the paper, redrawn: PPO (top) and GRPO (bottom). Based on Shao et al. (2024), Figure 4.

Reading it: read each half from the top down. In PPO (top) the policy writes one answer o. Two frozen models look at it: the reward model scores it, and the reference model supplies the KL fee, which is folded into the reward r (the ⊕). A second trained model, the value model, predicts v, and GAE combines r and v into advantages A. In GRPO (bottom) the value model is gone. The policy writes G answers to the same question, the reward model scores each, and a small calculation, each score minus the group's mean, divided by the group's standard deviation, turns scores into advantages. The reference model is still there, but its KL term moves out of the reward and into the loss. Solid boxes are the models being trained, dashed boxes the frozen ones: GRPO trains one large model instead of two.

The GRPO objective

Tiny example. The smallest possible group: G = 2 answers to one question, one right and one wrong, so (as the next subsection shows) their advantages are +1 and −1. The right answer has 2 tokens with ratios 1.1 and 1.3, the wrong one a single token with ratio 0.9. Take ε = 0.2, β = 0.04, and per-token KL estimates of 0.01, 0.02 and 0.01 (illustrative). The right answer's tokens contribute 1.1 − 0.0004 = 1.0996 and min(1.3, 1.2) − 0.0008 = 1.1992, averaging 1.1494; the wrong answer's token contributes −0.9 − 0.0004 = −0.9004. The objective is the average over the group: (1.1494 − 0.9004) / 2 = 0.1245.

In words: “for a question, sample a group of G answers from the old policy; for every token of every answer, take PPO's clipped ratio-weighted advantage minus β times a per-token estimate of the KL divergence from the reference model; average over each answer's tokens, then over the answers in the group.”

With the numbers: ((1.0996 + 1.1992) / 2 + (−0.9004)) / 2 = (1.1494 − 0.9004) / 2 = 0.1245.

In Python:

def clip(x, lo, hi):
    return max(lo, min(x, hi))
eps, beta = 0.2, 0.04
# each answer: its advantage Â_i, and per token (ratio, KL estimate)
group = [(+1.0, [(1.1, 0.01), (1.3, 0.02)]), (-1.0, [(0.9, 0.01)])]
def answer_term(A, tokens):
    # (1/|o_i|) Σ_t {min[r Â, clip(r) Â] − β D_KL}
    per_token = [min(r * A, clip(r, 1 - eps, 1 + eps) * A) - beta * kl for r, kl in tokens]
    return sum(per_token) / len(per_token)
[round(answer_term(A, toks), 4) for A, toks in group]  # → [1.1494, -0.9004]
# (1/G) Σ_i
round(sum(answer_term(A, toks) for A, toks in group) / len(group), 4)  # → 0.1245

Two design choices stand out. First, the advantage Âi,t comes only from comparing the answers in the group, which the paper notes fits how reward models are trained: on comparisons between answers to the same question. Second, the KL term sits in the loss rather than in the reward, so it does not complicate the advantage.

The KL estimator

Tiny example. At one token the policy gives probability 0.5 and the reference 0.4. The paper's estimate uses the ratio the other way up, reference over policy: 0.4 / 0.5 = 0.8, and computes 0.8 − ln 0.8 − 1 = 0.8 + 0.223 − 1 = 0.023. At a token where the policy gives 0.3 and the reference 0.4, the ratio is 1.333 and the estimate 1.333 − 0.288 − 1 = 0.046. Both are positive, even though the policy moved in opposite directions.

In words: “call x the reference's probability for the token divided by the policy's; the estimate is x minus the log of x minus 1, which is zero when the two agree and positive otherwise.”

With the numbers: 0.8 − ln 0.8 − 1 = 0.023; 1.333 − ln 1.333 − 1 = 0.046. Averaged over a whole distribution, it recovers the exact KL divergence: for a policy (0.5, 0.3, 0.2) and a reference (0.4, 0.4, 0.2), both come to 0.0253.

In Python:

import math
def k(pi_theta, pi_ref):
    # x − log x − 1, with x = π_ref / π_θ
    x = pi_ref / pi_theta
    return x - math.log(x) - 1
round(k(0.5, 0.4), 3), round(k(0.3, 0.4), 3)  # → (0.023, 0.046)
# unbiased: averaged over tokens drawn from π_θ, it equals the exact KL(π_θ ‖ π_ref)
pi_theta, pi_ref = [0.5, 0.3, 0.2], [0.4, 0.4, 0.2]
exact = sum(p * math.log(p / q) for p, q in zip(pi_theta, pi_ref))
average = sum(p * k(p, q) for p, q in zip(pi_theta, pi_ref))
round(exact, 4), round(average, 4)  # → (0.0253, 0.0253)

Hover or tap the chart to compare the two per-token estimates at any ratio.

Reading it: the x-axis is the per-token ratio x = πref / πθ, from 0.2 to 3; x = 1 means the policy and the reference agree on that token. The dashed line is the simplest per-token estimate, the log of πθ / πref (which is −log x): it is right on average but negative whenever the policy has made a token less likely than the reference did, so a single token can report a negative divergence. The solid line is the paper's estimate, x − log x − 1. It touches zero only at x = 1 and is positive everywhere else, which is what the paper means by “guaranteed to be positive”, and it is still right on average because the extra x − 1 averages to zero under the policy (the probabilities of the reference add to 1). The paper cites Schulman (2020) for this estimator.

Why it matters today

This subsection is the whole of GRPO: PPO's clip, a group baseline instead of a value network, and a per-token KL estimate in the loss. The lesson builds it in two pieces, group_advantages and clipped_policy_update, and runs it as train_grpo. The KL divergence itself is kl_divergence in the training stages lesson.

4.1.2 Outcome supervision RL with GRPO · original

Everyday picture

A teacher sets one question to a small group and marks on a curve: above the group's average is good, below is bad, and the size of the mark depends on how spread out the group was. “Outcome” supervision means only the final answer is scored, and every word of an answer shares its mark.

Tiny example

The running example: rewards 0.9, 0.1, 0.3 and 0.7 for the four answers to “What is 12 × 3?”. The mean is 0.5. The deviations are 0.4, −0.4, −0.2 and 0.2, whose squares average (0.16 + 0.16 + 0.04 + 0.04) / 4 = 0.1, so the standard deviation is √0.1 = 0.316. Dividing each deviation by it gives the advantages +1.26, −1.26, −0.63, +0.63. Every token of answer 1 gets +1.26.

In words: “every token of answer i gets the same advantage: how far the answer's reward sits above the group's average, measured in units of the group's spread.”

With the numbers: (0.9 − 0.5) / 0.316 = 1.26; (0.1 − 0.5) / 0.316 = −1.26; (0.3 − 0.5) / 0.316 = −0.63; (0.7 − 0.5) / 0.316 = 0.63.

In Python:

import statistics
r = [0.9, 0.1, 0.3, 0.7]
mean_r = statistics.mean(r)
std_r = statistics.pstdev(r)
round(mean_r, 3), round(std_r, 3)  # → (0.5, 0.316)
# r̃_i = (r_i − mean(r)) / std(r)
[round((r_i - mean_r) / std_r, 2) for r_i in r]  # → [1.26, -1.26, -0.63, 0.63]
# the paper doesn't say which standard deviation; with the sample one (G − 1) the sizes change, the signs don't
[round((r_i - mean_r) / statistics.stdev(r), 2) for r_i in r]  # → [1.1, -1.1, -0.55, 0.55]

Try it: drag the four rewards. Watch the advantages always add up to zero, and watch what happens when you make all four equal.

Reading it: each slider is one answer's reward; each bar is that answer's advantage, pointing right when it beats the group's average and striped when it falls below. Because each advantage is a distance from the group's own mean, they always add up to zero: GRPO never pushes every answer up at once, it moves probability from worse answers to better ones. Shift all four rewards up by the same amount and nothing changes, since only the differences matter. Scale the differences down and the advantages do not shrink, since dividing by the spread puts every group on the same scale. And make all four equal: the spread is zero, there is nothing to compare, and every advantage is zero, so that question teaches nothing this round (the widget, like the lesson's code, sets the advantages to zero rather than dividing by zero).

Why it matters today

The group mean is a baseline that costs no extra model, and the division by the spread makes easy and hard questions pull with similar force. That second part has a price: when a group's rewards barely differ, dividing by a tiny spread turns small differences into full-size advantages. The reasoning lesson shows this magnifying a small length penalty until it decides how long the model thinks.

4.1.3 Process supervision RL with GRPO · original

Everyday picture

A maths teacher who marks only the final answer cannot tell a student which step went wrong. One who marks every line can. Process supervision scores each reasoning step, and a token is credited with the scores of its own step and every step after it.

Tiny example

A group of two answers. Answer 1 has three steps, scored 0.9, 0.8 and 0.9 by a process reward model; answer 2 has two, scored 0.9 and 0.2 (its second step is where it went wrong). All five step scores are normalized together: mean 0.74, standard deviation 0.273, giving 0.59, 0.22, 0.59 for answer 1 and 0.59, −1.98 for answer 2. A token's advantage is the sum of the normalized scores from its step to the end: answer 1's steps get 1.39, 0.81, 0.59; answer 2's get −1.39, −1.98.

In words: “normalize every step's score against all the step scores in the group; then each token's advantage is the sum of the normalized scores of the step it sits in and all the steps after it.”

With the numbers: for answer 1, 0.59 + 0.22 + 0.59 = 1.39 in step 1, 0.22 + 0.59 = 0.81 in step 2, 0.59 in step 3; for answer 2, 0.59 − 1.98 = −1.39 in step 1 and −1.98 in step 2.

In Python:

import statistics
R = [[0.9, 0.8, 0.9], [0.9, 0.2]]
flat = [x for steps in R for x in steps]
mean_R, std_R = statistics.mean(flat), statistics.pstdev(flat)
round(mean_R, 2), round(std_R, 3)  # → (0.74, 0.273)
norm = [[(x - mean_R) / std_R for x in steps] for steps in R]
# Â for a token in step j: Σ over this step and every later one
[[round(sum(steps[j:]), 2) for j in range(len(steps))] for steps in norm]  # → [[1.39, 0.81, 0.59], [-1.39, -1.98]]

Try it: switch between the two kinds of supervision for the same two answers. Outcome supervision uses only whether each final answer was right (answer 1 right, answer 2 wrong).

Supervision:

Reading it: each row is one step of one answer, and its bar is the advantage every token in that step receives (striped when negative). Under outcome supervision each answer is one block: all of answer 1 gets +1 and all of answer 2 gets −1, so answer 2's perfectly good first step is punished as hard as its mistake. Under process supervision the credit is local: answer 1's advantage shrinks as fewer good steps lie ahead, and answer 2's mistake in step 2 gets the largest penalty, −1.98. Its first step is still negative (−1.39), because the sum looks ahead to the mistake it led to, but it is punished less than the mistake itself.

Why it matters today

In the paper's comparison (§5.2.1, Figure 5) process supervision beats outcome supervision. Process reward models are expensive to build, though, and the method that later became famous for training reasoning models, DeepSeek-R1, used outcome rewards from simple rule-based checks. The reasoning lesson compares outcome and process verifiers, with first_bad_step as a tiny step checker.

4.1.4 Iterative RL with GRPO · original

Everyday picture

As a student improves, a judge trained on beginner work stops being able to tell good from great. So every so often, retrain the judge on the student's recent work (keeping some older cases so it does not forget), and restart the leash from where the student now stands.

for iteration = 1 … I for step = 1 … M πref ← πθ sample a batch of questions πold ← πθ G answers per question rewards from rφ group-relative advantages μ GRPO iterations:maximize the objective retrain rφ on new samplesreplay: 10% historical data

Hover or tap a block. The outer dashed box is one iteration; the inner one is one training step.

Algorithm 1 of the paper, drawn as a diagram: iterative GRPO. Based on Shao et al. (2024), Algorithm 1.

Reading it: the outer dashed box is one iteration, the inner one a training step. Each iteration begins by freezing a copy of the current policy as the new reference, so the KL leash is re-anchored where the policy now stands. Inside, each step freezes the sampler (πold), samples G answers per question, scores them with the reward model rφ, turns scores into group-relative advantages and takes μ updates. After M steps the reward model itself is retrained on fresh samples from the improved policy, with 10% of historical data replayed so it keeps what it learned. With I = 1 and no retraining, this is plain GRPO.

Why it matters

A fixed learned reward model can be outgrown or gamed (reward hacking); refreshing it is one defence. In the paper's experiments (Figure 6) the iterations help, the first one most.

4.2 Training and evaluating DeepSeekMath-RL · original

The setup

  • RL starts from DeepSeekMath-Instruct 7B and uses only the chain-of-thought questions related to GSM8K and MATH from the SFT data, about 144K questions. Every other benchmark is kept out of RL on purpose, to see whether gains spread to tasks never trained on.
  • The reward model is trained from DeepSeekMath-Base 7B (learning rate 2 × 10−5), with training data built following Wang et al. (2023b).
  • GRPO: policy learning rate 1 × 10−6, KL coefficient β = 0.04, 64 sampled answers per question, maximum length 1024 tokens, batch size 1024, and a single policy update after each round of sampling.

That last point matters for reading equation (3): with one update per sample, πθ equals πold when the gradient is taken, every ratio is 1, and the clip never activates. The appendix uses exactly this simplification.

Selected rows of Table 5, reproduced with attribution (Shao et al., 2024). Top-1 accuracy with a chain of thought; Instruct and RL are the DeepSeekMath 7B models
ModelGSM8KMATHMGSM-zhCMATH
GPT-492.0%52.9%n/a86.0%
Gemini Ultra94.4%53.2%n/an/a
Instruct82.9%46.8%73.2%84.6%
RL88.2%51.7%79.6%88.8%

Reading it: each pair of bars is one benchmark: the solid bar is DeepSeekMath-Instruct 7B before RL, the striped bar the same model after GRPO, on a scale from 40% to 100% so the gaps are visible. GRPO lifts all four. GSM8K and MATH are in-domain (RL trained on questions like them); MGSM-zh and CMATH are Chinese benchmarks it never trained on, and they improve too, by 6.4 and 4.2 points. The paper reports the same pattern with tool use, and says the RL model improves over the Instruct model on every benchmark in Table 5.

Why it matters

About 144K questions and a policy learning rate of 10−6 bought roughly 5 points on both in-domain benchmarks, on top of a model that was already strong, without a value network.

5 Discussion · original

5.1 Lessons learnt in pre-training · original

Everyday picture

Does learning to program make you better at maths? Does reading research papers? The paper tests both with small models (these experiments use an earlier, 89B-token version of the corpus).

5.1.1 Code training benefits mathematical reasoning

Selected rows of Table 6, reproduced with attribution (Shao et al., 2024). DeepSeek-LLM 1.3B
TrainingGSM8KMATHGSM8K, PythonMATH, Python
400B general, then 150B math19.1%14.4%14.3%6.7%
400B code, then 150B math21.9%15.3%17.4%9.4%
150B math only20.5%13.1%11.4%6.5%
400B code and 150B math, mixed17.6%12.1%19.7%13.5%

Code first, then maths, gives the best maths without tools; mixing code and maths in one stage gives the best maths with Python and avoids the catastrophic forgetting of coding that two stages cause (Table 7), but hurts maths without tools, which the paper guesses is too much for a 1.3B model to absorb at once. This is why DeepSeekMath starts from a code model.

5.1.2 ArXiv papers seem ineffective

Training on arXiv-only corpora (MathPile, and ArXiv-RedPajama at 28.0B tokens) brought no notable improvement, and sometimes a drop, on every maths benchmark tried, for both a 1.3B and a 7B model. For the 7B model, MATH went from 12.5% with no maths training to 11.5% (MathPile) and 11.1% (ArXiv-RedPajama) (Table 8). The paper is careful to limit the claim: it did not test arXiv data on other tasks such as turning formal proofs into informal ones, mixed with other data, or at larger scale.

Why it matters

“More text about maths” and “text that teaches problem solving” are different things; the pretraining lesson shows how a data mixture is chosen with exactly this kind of small ablation.

5.2.1 Towards a unified paradigm · original

“Within this paradigm, all methods are conceptualized as either direct or simplified RL techniques.”Shao et al. (2024), §5.2.3

Everyday picture

Every training method for a language model does the same thing at every token: it nudges the probability of that token up or down. What differs is only which tokens it looks at, and how hard it pushes each one. The paper calls the push the gradient coefficient.

Tiny example

At one token the model's vocabulary is three tokens with probabilities 0.5, 0.3 and 0.2, and the answer used the first. The direction that makes that token more likely is “1 for it, minus every probability”: (0.5, −0.3, −0.2) in logit space. GRPO scales it by answer 1's advantage, 1.26, giving (0.63, −0.38, −0.25). Rejection sampling fine-tuning would scale it by 1 (answer 1 is right); for answer 2, a wrong answer, RFT scales it by 0, while GRPO scales it by −1.26, actively pushing that token down.

In words: “for any method A: take question-answer pairs from some data source, and for every token add the direction that makes that token more likely, scaled by the method's gradient coefficient, which depends on the reward signal.”

With the numbers: ∇ log π for the chosen token = (1 − 0.5, −0.3, −0.2) = (0.5, −0.3, −0.2); × 1.26 = (0.63, −0.38, −0.25) under GRPO for answer 1; × 0 under RFT for a wrong answer.

In Python:

pi = [0.5, 0.3, 0.2]
chosen = 0
# ∇_θ log π_θ(o_t) for a softmax over logits: 1[k = chosen] − π(k)
grad_log_pi = [(k == chosen) - pi[k] for k in range(3)]
grad_log_pi  # → [0.5, -0.3, -0.2]
def contribution(GC):
    return [round(GC * g, 2) for g in grad_log_pi]
# GRPO, answer 1: GC = its advantage
contribution(1.26)  # → [0.63, -0.38, -0.25]
# GRPO, answer 2 (wrong): pushed down
contribution(-1.26)  # → [-0.63, 0.38, 0.25]

The paper fills in the three components for six methods (Table 10). “Online” means the answers are sampled from the model being trained, as it changes; “offline” means they were all sampled once, from the SFT model, before training began.

Table 10 of the paper, in words (Shao et al., 2024)
MethodData sourceRewardGradient coefficient
SFThuman-selected question-answer pairsnone1
RFTanswers sampled once from the SFT model (offline)rule: is the answer correct?1 if correct, 0 if not
DPOpairs sampled once from the SFT model (offline)ruleσ of the reward gap (below)
Online RFTanswers sampled from the current policy (online)rule1 if correct, 0 if not
PPOanswers sampled from the current policy (online)reward modelthe advantage At
GRPOgroups of answers sampled from the current policy (online)reward modelgroup advantage plus a KL term

In words: “RFT pushes every token of a correct answer up by the same amount and ignores wrong ones; PPO pushes by the advantage; DPO pushes the preferred answer up and the rejected one down, hardest when the model currently favours the wrong one; GRPO pushes by the group advantage, plus a small correction pulling each token's probability back towards the reference model.”

With the numbers: for the running group, RFT gives (1, 0, 0, 1); GRPO at the start of training (every πref / πθ = 1) gives (1.26, −1.26, −0.63, 0.63). If training has made a token of answer 2 more likely than the reference did (πref / πθ = 0.8), GRPO's push on it becomes −1.265 + 0.04 × (0.8 − 1) = −1.273 (using the unrounded advantage). DPO at the start, with πθ = πref, gives σ(0) = 0.5.

In Python:

import math, statistics
correct = [1, 0, 0, 1]
r = [0.9, 0.1, 0.3, 0.7]
beta = 0.04
# RFT: I(o), 1 for a correct answer
correct  # → [1, 0, 0, 1]
mean_r, std_r = statistics.mean(r), statistics.pstdev(r)
A = [(r_i - mean_r) / std_r for r_i in r]
def gc_grpo(A_i, ref_over_theta=1.0):
    return A_i + beta * (ref_over_theta - 1)
[round(gc_grpo(a), 2) for a in A]  # → [1.26, -1.26, -0.63, 0.63]
round(gc_grpo(A[1], 0.8), 3)  # → -1.273
def sigma(z):
    return 1 / (1 + math.exp(-z))
# DPO when the policy still equals the reference: both log ratios are 0
sigma(beta * 0.0 - beta * 0.0)  # → 0.5

Reading it: each pair of bars is one answer in the running group; the solid bar is its gradient coefficient under online RFT, the striped bar under GRPO. RFT's bars are all at 1 or 0: it treats the two right answers identically and simply skips the wrong ones. GRPO's bars have sizes and signs: the better-scored right answer gets the bigger push, and the wrong answers are pushed down, the one the reward model rated worse more strongly. That is the paper's “observation about gradient coefficient”: GRPO beats online RFT in its experiments because it can reinforce and penalize in proportion.

What the experiments show

  • Data source. Online RFT beats offline RFT, and the gap opens late in training (Figure 5, on DeepSeekMath-Instruct 1.3B): early on the policy still resembles the SFT model, so their samples differ little; later, fresh samples from the current policy are worth more.
  • Gradient coefficient. GRPO beats online RFT, and GRPO with process supervision beats GRPO with outcome supervision: fine-grained, step-aware coefficients help.
  • Iterative RL. Two rounds of iterative GRPO (Figure 6, on the 7B model) improve on plain GRPO, the first round most.

Why it matters today

This single formula is a good way to hold the whole zoo of post-training methods in your head. The training stages lesson builds SFT and DPO (dpo_loss), and the reinforcement learning lesson builds REINFORCE, PPO and GRPO: read side by side, they differ in exactly the coefficient this table lists.

5.2.2 Why RL works? · original

“As shown in Figure 7, RL enhances Maj@K’s performance but not Pass@K.”Shao et al. (2024), §5.2.2

Everyday picture

A student who sometimes gets a question right and sometimes wrong can be trained to be right reliably. That is different from teaching them to answer questions they could never answer at all. The paper's evidence suggests RL here mostly did the first.

Tiny example

Pass@K asks whether any of K sampled answers is right; Maj@K asks whether the most common answer among K is right. Take a question the model answers correctly 40% of the time. With 64 samples it will almost surely produce a right answer somewhere (pass@64 ≈ 1), but the wrong answer usually wins the vote. If training lifts that 40% to 55%, pass@64 is still about 1, while the majority flips to the right answer. (This is a simplified picture in which all wrong answers agree with each other; the percentages are illustrative.)

Hover or tap the chart to read all four curves at any K.

Reading it: this is the shape of the paper's Figure 7, drawn for six illustrative questions whose chance of a right answer is 0.9, 0.7, 0.55, 0.4, 0.2 and 0.05 before RL and 0.97, 0.9, 0.8, 0.55, 0.25 and 0.05 after (the values are illustrative, not the paper's). The x-axis is the number of samples K, doubling at each step from 1 to 64. At K = 1 both kinds of score are just the top-1 accuracy, and RL is ahead. As K grows the two Pass@K curves (for any right answer) climb together and meet near 1 by K = 32: every question RL helps was already solvable sometimes, so sampling enough answers finds a right one either way. The two Maj@K curves stay apart: RL's majority is right more often at every K. That is the paper's finding, and its reading of it: RL made the output distribution more robust, boosting right answers that were already in the model's top K, rather than adding new capability.

Why it matters today

Keep this in mind whenever RL results are reported: did training teach new reasoning, or make the model more reliable at what it could already sometimes do? Comparing Pass@K at large K before and after is one way to tell. The functions behind these curves are in the reasoning lesson: pass_at_n and majority_accuracy.

5.2.3 How to achieve more effective RL? · original

Everyday picture

Every method in equation (5) has three dials: where the practice questions and answers come from, how the push is computed, and who does the grading. The paper suggests a direction for each.

  • Data source. RL here used only the SFT questions and plain nucleus sampling, which the authors suspect is why only Maj@K improved. They propose out-of-distribution questions, smarter search-based sampling (tree search), and faster inference, since exploration speed limits RL.
  • Algorithms. Every method trusts its reward completely, but rewards are noisy: even the carefully annotated PRM800K dataset of step labels contains about 20% incorrect annotations, the paper notes. They call for algorithms robust to noisy rewards, in the spirit of weak-to-strong learning.
  • Reward function. Reward models that generalize to new questions and decoding methods, that express their uncertainty, and cheaper ways to build high-quality process reward models.

Why it matters today

Read with hindsight, the most consequential suggestion was implicit: when the reward is a rule rather than a learned model, there is no reward model to generalize badly or be gamed. The verifiable reward section of the reinforcement learning lesson runs GRPO against exactly such a checker, verify.

6 Conclusion, limitation, and future work · original

“Our extensive ablation study shows web pages offer significant potential for high-quality mathematical data, while arXiv may not as beneficial as we expected.”Shao et al. (2024), §6

What the paper concludes

An open 7B model, continued-pretrained for 500B tokens (120B of them maths from Common Crawl), then instruction-tuned and trained with GRPO, outperforms every open model on MATH and approaches the closed ones. GRPO improves maths with less memory than PPO, and works even on a model that already scores highly.

What it does not do

  • Geometry and theorem proving are weaker than in closed models; in a dry run the model could not handle problems about triangles and ellipses, which the authors suspect reflects bias in data selection.
  • Unlike GPT-4, it does not improve when given few-shot examples: its zero-shot and few-shot scores are similar, which the authors attribute to model scale.

Why it matters

The limitations are about the model; the method's influence went well beyond maths.

Appendix A.1: where each gradient coefficient comes from · original

Everyday picture

The appendix is the proof behind the table in §5.2.1: for each method it writes the objective, takes its gradient, and reads off the coefficient in front of ∇ log π. The step worth seeing in full is the one that turns GRPO's KL estimate into the small correction β(πref / πθ − 1).

Tiny example

Take one token with πθ = 0.5 and πref = 0.4, so x = πref / πθ = 0.8. The KL estimate is x − log x − 1. Nudge log πθ up by a tiny 0.001: x shrinks by a factor e−0.001 and the estimate rises by about 0.0002, a slope of 1 − x = 0.2 with respect to log πθ. With the −β in front, the KL term's slope is −β(1 − x) = β(x − 1) = 0.04 × (−0.2) = −0.008, which is exactly GRPO's correction term.

In words: “x depends on θ only through πθ in its denominator, so raising log πθ by a little lowers x in proportion to x; the estimate x − log x − 1 then changes by (1 − x) per unit of log πθ, and the −β in front turns that into β(x − 1).”

With the numbers: x = 0.8, so β(x − 1) = 0.04 × (−0.2) = −0.008 per unit of ∇ log πθ: a gentle push down on a token the policy already favours more than the reference does.

In Python:

import math
beta, pi_ref = 0.04, 0.4
def kl_term(log_pi_theta):
    # −β (x − log x − 1), with x = π_ref / π_θ
    x = pi_ref / math.exp(log_pi_theta)
    return -beta * (x - math.log(x) - 1)
log_pi, h = math.log(0.5), 1e-6
# the slope, measured by nudging log π_θ
round((kl_term(log_pi + h) - kl_term(log_pi - h)) / (2 * h), 4)  # → -0.008
# the formula: β (x − 1)
x = pi_ref / 0.5
round(beta * (x - 1), 4)  # → -0.008

The appendix also derives the other rows. For SFT the coefficient is 1 because the objective is just the average log-probability of human-written answers. For RFT it is the indicator 𝕀(o), because incorrect samples are dropped. For PPO, assuming a single update per round so that πold = πθ, the min and the clip disappear and the coefficient is the advantage At. GRPO adds the KL correction above to its group advantage.

Why it matters

The derivation shows why the KL term is gentle: its push on any token is β times how far the policy's and the reference's probabilities have drifted apart, 0.008 in the example against a group advantage of about 1.

What changed since 2024

In the paperWhat followedRead more
GRPO with a learned reward model, on mathsDeepSeek-R1 (2025) used GRPO with rule-based rewards (a correct final answer and a required format) and reported long, self-checking chains of thought emerging from RL alonereasoning lesson
Advantages divided by the group's standard deviationLiu et al. (2025) analyse biases in GRPO's normalization, including the effect on answer lengthreasoning lesson
PPO with a value network, the standard for RLHFA group baseline, as in GRPO, is an alternative wherever several answers per prompt are affordablePPO companion
Maj@K improves, Pass@K does notAn open question whether RL with verifiable rewards adds reasoning ability or sharpens what is there§5.2.2 above

Glossary

Every term with hover guidance on this page, in one place.