An annotated companion · AI Primer

DeepSeek-R1, annotated

About this page. This is a companion, not a copy. It follows the second arXiv version of the paper (January 2026), which reorganised the January 2025 report into a main paper plus supplementary sections A to J; every section heading links to that version. The paper is posted under arXiv's standard licence, not an open one, so this page quotes at most a sentence or two per section (clearly marked), shows only a few selected numbers from its tables with attribution, and redraws its diagrams from scratch. The equations are reproduced with every symbol decoded. Numbers marked illustrative are ours, chosen to be checked by hand.

How to read this page

  • Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
  • The diagrams are live: hover or tap a part. Four pieces respond to you: a group of answers you can mark right or wrong, the leash's estimator, a toy that learns how long to think, and the paper's scores at each training stage.

Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The algorithm underneath is a policy gradient; if REINFORCE, baselines or PPO's clipping are new, the REINFORCE companion and the reinforcement learning lesson build them from scratch, and the reasoning models lesson trains a toy version of this paper's recipe.

Abstract · original

“Here we show that the reasoning abilities of LLMs can be incentivized through pure reinforcement learning (RL), obviating the need for human-labeled reasoning trajectories.”DeepSeek-AI (2025), Abstract

Everyday picture

A maths workbook with only the answers printed in the back. Nobody shows you a single worked solution. You attempt each problem several ways, check the back, and do more of whatever worked. Over thousands of problems you discover good habits yourself: write things down, double-check the tricky step, try another approach when stuck. This paper does that to a base model, at scale.

What the paper claims

  • Reinforcement learning alone, rewarding only correct final answers, turns a base model into a strong reasoner. The resulting model, DeepSeek-R1-Zero, was never shown a human-written chain of thought.
  • Reasoning behaviours such as self-reflection, verification and switching strategies emerge on their own, and responses grow steadily longer during training.
  • A multi-stage pipeline, DeepSeek-R1, keeps that reasoning while fixing readability and adding general helpfulness and safety.
  • The reasoning transfers to smaller models by distillation: fine-tuning them on R1's outputs.

Why it matters today

It made an ingredient list public: a strong base model, hard problems with checkable answers, a group-based policy gradient, and a lot of compute. That recipe, with variations, is how open reasoning models have been trained since.

1 Introduction · original

Everyday picture

You can teach a child long division by making them copy your worked examples, or by giving them problems and an answer key. Copying is fast but caps them at your own method, habits and mistakes included. The answer key is slower but lets them find methods you never showed them.

The argument

  • Reasoning in language models came from scale and from chain-of-thought prompting, then from training on human-written reasoning traces.
  • Human traces are expensive, carry human biases, and cap the model at human-style reasoning.
  • So: start from DeepSeek-V3-Base, use GRPO as the learning algorithm, reward only whether the final answer matches the ground truth, and skip the usual supervised fine-tuning stage before RL entirely.

The result, R1-Zero, reasons well but writes hard-to-read chains that can mix English and Chinese, and it is narrow. Section 3's R1 adds a little human-shaped data back in to fix that. Section F's distilled small models make it cheap.

Why it matters

The bet is that the base model already contains the ability, and RL only has to find and reinforce it. The conclusion returns to this: what unlocks reasoning is “hard reasoning questions, a reliable verifier, and sufficient computational resources”, not large-scale human annotation.

2 DeepSeek-R1-Zero · original

Everyday picture

A student with the answer key and nothing else. For each problem they write sixteen attempts, check them all, and shift towards whatever the right attempts did. The rest of this section is that sentence made precise: how attempts are compared (2.1), what counts as right (2.2), and what the student turns into (2.3).

2.1 Group Relative Policy Optimization · original

Everyday picture

A teacher doesn't need a separate examiner to predict how hard each question is. Give the same question to a group of students and grade each answer against the others: on a question everyone gets right, getting it right earns little; on a hard one, the lone right answer stands out. GRPO grades a model's answers the same way, against other answers to the same question.

Tiny example

One question, a group of G = 4 answers. Two are right (reward 1) and two wrong (reward 0). The group averages 0.5 with a spread of 0.5, so the right answers get advantage +1 and the wrong ones −1. Since the batch was sampled, the model has already moved a little: it now gives the four answers probability ratios of 1.3, 0.9, 1.05 and 0.7 compared with when they were sampled. With the clip range ε = 0.2, the first answer's credit is capped at 1.2 (it has already been pushed up enough), while the fourth answer, a right one that has become less likely, counts in full.

PPO q Policytrained o Reference Reward Valuetrained reward− KL penalty GAE → A GRPO KL added straight to the loss q Policytrained o1o2⋮oG Reference Reward r1⋮rG group mean and spread→ A1 … AG

Hover or tap a part. Compare the top row (PPO) with the bottom row (GRPO): which boxes disappear?

The paper's Figure 3 (Supplementary A.3), redrawn: PPO against GRPO. Based on DeepSeek-AI (2025), Figure 3.

Reading it: both rows start from a question q and the policy being trained, and both end in an advantage A, a number saying how much better than expected each output was. In PPO (top), one output o is scored by a reward model and, separately, by a value model that must itself be trained and is about as big as the policy; the KL leash is folded into the reward, and a procedure called GAE combines reward and value into the advantage. In GRPO (bottom) the value model is gone. Instead the policy writes a group of G outputs, each is scored, and the group's own mean and spread turn scores into advantages. The KL leash, instead of being part of the reward, goes straight into the loss (the long wire back to the policy).

The math: the GRPO objective

In words: “for each question, sample a group of G answers from the model as it was; for each answer, multiply its advantage by how much more likely the current model makes it, but stop counting once that ratio leaves 1 ± ε in the direction the advantage wants, and always take the more pessimistic of the two; average over the group, subtract β times the distance from the reference model, and push the model's weights to make the result bigger.”

With the numbers (illustrative): advantages (+1, −1, −1, +1), ratios (1.3, 0.9, 1.05, 0.7), ε = 0.2. The four clipped terms are min(1.3, 1.2) = 1.2, min(−0.9, −0.9) = −0.9, min(−1.05, −1.05) = −1.05 and min(0.7, 0.8) = 0.7. Their average is −0.0125. With β = 0.001, the leash term barely moves it.

In Python:

def clip(x, lo, hi):
    return max(lo, min(x, hi))
A = [1, -1, -1, 1]
# ρ_i = π_θ(o_i|q) / π_θold(o_i|q) for each answer, illustrative
rho = [1.3, 0.9, 1.05, 0.7]
eps = 0.2
# min(ρ A, clip(ρ, 1 − ε, 1 + ε) A) for each answer
terms = [min(r * a, clip(r, 1 - eps, 1 + eps) * a) for r, a in zip(rho, A)]
[round(x, 2) for x in terms]  # → [1.2, -0.9, -1.05, 0.7]
# (1/G) Σ_i over the group of G = 4
round(sum(terms) / len(terms), 4)  # → -0.0125

The paper writes each output's ratio as one number for the whole answer. The DeepSeekMath paper, where GRPO was introduced, writes the same objective token by token: the ratio and the leash are computed at every token and averaged over each answer's tokens, and with a reward only at the end, every token of an answer shares that answer's advantage.

The advantage, from the group alone

In words: “an answer's advantage is how far its reward sits above its group's average, in units of the group's spread.”

With the numbers: rewards (1, 0, 0, 1): mean 0.5, spread 0.5, advantages (+1, −1, −1, +1). R1-Zero samples G = 16: if 4 of 16 are right, the mean is 0.25 and the spread 0.433, so each right answer gets +1.73 and each wrong one −0.58. If all 16 are right (or all wrong), every advantage is 0.

In Python:

import statistics
def advantages(r):
    mu, sd = statistics.mean(r), statistics.pstdev(r)
    # a group that all agree carries no signal: every advantage is 0
    return [round((r_i - mu) / sd, 2) if sd else 0.0 for r_i in r]
advantages([1, 0, 0, 1])  # → [1.0, -1.0, -1.0, 1.0]
# G = 16 with 4 right: the right ones, then one of the wrong ones
a = advantages([1] * 4 + [0] * 12)
a[0], a[-1]  # → (1.73, -0.58)
advantages([1] * 16)[:3]  # → [0.0, 0.0, 0.0]

Try it: tap answers to mark them right or wrong, and watch every answer's advantage change. Then make them all the same.

Reading it: each chip is one of the 16 answers R1-Zero samples per question, with its advantage underneath. Filled chips marked ✓ are right (reward 1); outlined chips marked ✗ are wrong (reward 0). Start from 4 right: each right answer gets +1.73, each wrong one −0.58, and the advantages always add up to zero, so the group pushes as much down as up. Mark more answers right and each right answer's advantage shrinks: success is becoming expected. Mark all 16 right, or all wrong, and every advantage drops to 0: the question has nothing left to teach, which is why training needs questions the model solves only sometimes. The advantages here are exact, computed as in the Python above.

The settings for R1-Zero

  • Learning rate 3 × 10−6, KL coefficient β = 0.001, sampling temperature 1.
  • 16 outputs per question, each at most 32,768 tokens, raised to 65,536 after step 8,200; both accuracy and response length jump at that step.
  • 32 questions per step, so 512 outputs per step; 10,400 steps in all, 1.6 passes over the data.
  • Each round of generation writes 8,192 outputs, split into 16 mini-batches and used for one pass only; every 400 steps the reference model is replaced by the latest policy.

Why it matters

Dropping the value model halves the number of big models that must be trained, and the paper argues (Supplementary A.3) that a value model is especially hard to train for long chains of thought, where a chain can later revise or contradict what it wrote earlier. The lesson's group_advantages and train_grpo build this, and ppo_clipped_objective is the clipped term.

The leash, and its estimator · original

Everyday picture

A dog on a long leash can explore, but not run off. The leash ties the policy to a reference model, and its strength is β. What it measures is how far the policy's probabilities have drifted from the reference's: the KL divergence. The paper estimates that distance from the samples it already has, using a formula that can never come out negative.

Tiny example

The policy now gives an answer probability 0.6; the reference gave it 0.3. The ratio reference over policy is 0.5, and the estimate is 0.5 − ln 0.5 − 1 = 0.5 + 0.693 − 1 = 0.193. For an answer the policy has made less likely than the reference (0.4 against 0.7), the ratio is 1.75 and the estimate is 1.75 − ln 1.75 − 1 = 0.190. Both positive.

In words: “take the ratio of the reference's probability to the policy's for the answer just sampled; the estimate is that ratio, minus its logarithm, minus one.” The paper calls it an unbiased estimator of the KL divergence (after Schulman, 2020): averaged over answers the policy samples, it equals the true divergence.

With the numbers: 0.5 − ln 0.5 − 1 = 0.193. For a two-answer policy (0.6, 0.4) and reference (0.3, 0.7), the average of the estimate over the policy's own samples is 0.6 × 0.193 + 0.4 × 0.190 = 0.192, and the exact KL divergence is 0.6 ln(0.6/0.3) + 0.4 ln(0.4/0.7) = 0.192. They match.

In Python:

import math
def k3(pi, pi_ref):
    # π_ref/π_θ − log(π_ref/π_θ) − 1
    x = pi_ref / pi
    return x - math.log(x) - 1
round(k3(0.6, 0.3), 3), round(k3(0.4, 0.7), 3)  # → (0.193, 0.19)
policy, reference = [0.6, 0.4], [0.3, 0.7]
# the estimate averaged over the policy's own samples
round(sum(p * k3(p, r) for p, r in zip(policy, reference)), 3)  # → 0.192
# the exact KL divergence
round(sum(p * math.log(p / r) for p, r in zip(policy, reference)), 3)  # → 0.192

Try it: hover the chart, or tab to it and use the arrow keys, to compare this estimate with the simplest one, the log of the policy-over-reference ratio.

Hover the chart, or tab to it and use the arrow keys, to read both estimates at a ratio.

Reading it: the x-axis is the ratio πref/πθ for one sampled answer: 1 means the policy and reference agree on it, below 1 means the policy now favours it, above 1 means the policy has turned against it. The dashed line is the simplest single-sample estimate, log(πθ/πref), which is negative whenever the policy has made the answer less likely, so single samples can cancel out and hide drift. The solid line, the paper's estimate, is a bowl that touches zero only at a ratio of 1 and is positive everywhere else: every sample counts drift in either direction as cost. Both lines average to the same true divergence, but the solid one never rewards a sample for drifting.

Why it matters

Where the leash goes is a real design choice. PPO for chat models subtracts a KL penalty from the reward at every token, which, the paper argues, may penalise long answers simply for being long. GRPO puts the estimate straight into the loss instead, and R1-Zero also resets the reference to the current policy every 400 steps, so the leash limits each stretch of training without anchoring a model that has to travel far. The lesson's kl_penalised_reward is the reward-side version.

2.2 Reward design · original

“The reward is the source of the training signal, which decides the direction of RL optimization.”DeepSeek-AI (2025), §2.2

Everyday picture

An exam marked by a machine that only checks the boxed final answer, plus a rule that working must be written in the space provided. Nobody reads the working for quality. That is R1-Zero's whole reward.

Tiny example (illustrative scores)

The template asks the model to put its reasoning between <think> and </think> and its answer between <answer> tags. Take a problem whose answer is 7, and suppose each reward is 1 or 0:

Illustrative: four responses to a problem whose answer is 7, scored by the two rules.
ResponseAccuracyFormatTotal
reasoning in think tags, answer 7112
reasoning in think tags, answer 8011
no tags, answer 7101
no tags, answer 8000

In words: “the rule-based reward is the accuracy reward plus the format reward, weighted equally.”

With the numbers (illustrative): the four responses score 2, 1, 1 and 0.

In Python:

import re
def reward_rule(response, truth):
    formatted = bool(re.fullmatch(r"<think>.*</think>\s*<answer>.*</answer>", response, re.S))
    answer = re.search(r"<answer>(.*)</answer>", response)
    given = answer.group(1).strip() if answer else response.strip().split()[-1]
    # Reward_acc + Reward_format, 1 or 0 each here (illustrative)
    return int(given == truth) + int(formatted)
reward_rule("<think>3 + 4 = 7</think> <answer>7</answer>", "7")  # → 2
reward_rule("<think>3 + 4 = 8</think> <answer>8</answer>", "7")  # → 1
reward_rule("the answer is 7", "7")  # → 1

Accuracy is checked by rules: for maths, the final answer in a specified format (such as a box) is compared with the reference; for code, a compiler runs the answer against test cases. The paper deliberately uses no neural reward model, neither one that scores outcomes nor one that scores each step, for reasoning tasks. Its reason: such models are open to reward hacking during large-scale RL, and retraining them costs compute and complicates the pipeline.

Why it matters

This is a verifiable reward, and choosing one is the paper's quiet foundation: a check with no learned blind spots can be optimised against for thousands of steps without the policy learning to fool it. The lesson's verify is the simplest possible accuracy reward, and outcome_reward checks only a chain's final answer.

2.3 Incentivize reasoning capability in LLMs · original

Everyday picture

Nobody told the workbook student to show more working or to double-check. But attempts with more working and a double-check got the answer right more often, so, attempt by attempt, those habits won out.

What happened to R1-Zero

  • Accuracy. On the AIME 2024 maths competition, average pass@1 rose from 15.6% to 77.9% over training, and 86.7% with self-consistency (a majority vote over samples), above the average human competitor.
  • Length. The average response length on training problems grew steadily throughout, to hundreds and thousands of tokens. Nothing in the reward asked for it.
  • Behaviour. Reflection and systematic exploration of alternatives grew with the length; Supplementary C counts it.
“Wait, wait. Wait. That's an aha moment I can flag here.”An intermediate DeepSeek-R1-Zero, in the paper's Table 2

That line is the paper's aha moment: midway through squaring both sides of an equation, the model stops, says so, and starts re-evaluating step by step. The paper ties it to a sudden rise in the use of the word “wait” during training.

Tiny example: why length is rewarded without being asked for

The reasoning lesson's toy (our numbers, not the paper's): a problem needs d steps; each step slips 10% of the time, and a chain with room for two tries per step can catch a slip. An 8-step problem solved with one try per step (8 tokens) succeeds 0.98 = 43% of the time; with room for a recheck (16 tokens), 0.998 = 92%. Rewarded only for right answers, the model drifts to 16.

Try it: the toy below trains a tiny policy with GRPO to choose how long to think, for problems needing 1, 2, 4 or 8 steps. Start with no penalty, then charge for every token, then switch off GRPO's division by the group's spread.

Reading it: the x-axis is training iterations; the y-axis is the policy's average thinking length in tokens, averaged over the four problem sizes. The dashed line is always the correctness-only run; the solid line follows your settings (with no penalty and the box ticked, they coincide). Rewarded only for right answers, the length climbs from 2 tokens to about 7.8 and accuracy from 0.41 to 0.96: exactly R1-Zero's story in miniature, a longer chain because checking pays. Add a penalty of 0.01 per token and the length settles lower, and with the spread division on, the easiest problems collapse to 1 token and accuracy stops at 0.92: in a group where every answer is right, rewards differ only by the tiny penalty, and dividing by the tiny spread inflates it into a full ±1 advantage. Untick the box and the easiest problems keep 2 tokens (accuracy 0.93). The readout gives each run's numbers. Every run is seeded, so it repeats exactly.

Why it matters

This is the paper's headline: rather than being taught how to reason, the model was given the right incentive and found long, self-checking reasoning itself. The toy is built in the lesson as train_reasoner with success_probability, and the lesson's section on the price of thinking explains the penalty's surprising strength.

3 DeepSeek-R1 · original

Everyday picture

The self-taught student is brilliant but writes in a private shorthand, switches languages mid-sentence, and has only ever done maths and code. Before letting them tutor others, you show them a few neatly written solutions, let them practise again with a rule about staying in one language, then broaden their training to writing and conversation, and finally polish helpfulness and safety.

Tiny example: the stages

  1. Cold start. Thousands of long, readable, first-person reasoning examples fine-tune the base model into R1 Dev1.
  2. Reasoning RL. GRPO with rule-based rewards plus a language-consistency reward gives Dev2.
  3. Rejection sampling and SFT. About 600,000 correct reasoning samples from Dev2 and about 200,000 non-reasoning examples fine-tune the base model again: Dev3.
  4. RL for all scenarios. Rule-based rewards for reasoning plus reward models for helpfulness and harmlessness give DeepSeek-R1.
V3-BaseV3-BaseV3-Base RL: accuracyand format R1-Zero sample filter andrewrite cold startthousands SFT oncold start R1 Dev1 RL: rules andone language R1 Dev2 600K reasoning200K general SFT R1 Dev3 RL: rules andpreferences DeepSeek-R1

Hover or tap a box. Read the columns left to right, each top to bottom.

The paper's Figure 2, redrawn: the multi-stage pipeline of DeepSeek-R1. Based on DeepSeek-AI (2025), Figure 2.

Reading it: each column starts from the same base model, DeepSeek-V3-Base; only the data and the training change. The left column is R1-Zero, trained by RL alone, whose outputs are sampled, filtered for right answers and readable format, and rewritten into the cold-start data. The middle column fine-tunes a fresh base model on that data (Dev1) and runs reasoning RL with a language-consistency reward (Dev2). The right column samples Dev2 many times, keeps the right answers, adds general data from DeepSeek-V3's own fine-tuning set, fine-tunes a fresh base model on all of it (Dev3), and finishes with RL that mixes rule-based and preference rewards. Notice that R1-Zero is not an ancestor of R1's weights: it contributes only data, and so does Dev2.

Why it matters

This alternation, RL to discover, rejection sampling to harvest what was discovered, SFT to make it the new starting point, is the pattern most reasoning pipelines now follow in some form. The paper's Supplementary G states why both halves are needed: RL finds reasoning that human traces can't teach, and SFT covers tasks where no reliable reward exists.

3.1 Model-based rewards · original

Everyday picture

No program can check whether a poem is good or a reply is kind. For those, the paper hires a trained judge, a reward model, with two rules: the helpfulness judge reads only the final summary, so it cannot meddle with the reasoning, and the safety judge reads everything, reasoning included, because harm can hide in either.

Tiny example (illustrative)

To build training pairs for the helpfulness judge, DeepSeek-V3 compares two responses four times, each time assigning at random which one is shown as A and which as B, to cancel any preference for position; the four verdicts are averaged. Suppose the verdicts favour response A by 2, 1, 2 and 1 points: the average is 1.5. Only pairs whose score difference exceeds 1 are kept, so this pair is kept; a pair averaging 0.5 would be dropped as too close to call.

In words: “the helpfulness reward comes from a model trained on pairs of responses, judging which is better; the safety reward comes from a model that judges a single response as safe or unsafe.”

With the numbers (illustrative): verdicts (2, 1, 2, 1) average 1.5 > 1: kept. Verdicts (1, 0, 1, 0) average 0.5: dropped.

In Python:

# four judgments of one pair, positions assigned at random each time (illustrative scores)
def keep_pair(verdicts):
    delta = sum(verdicts) / len(verdicts)
    # keep only clear preferences: a score difference above 1
    return delta, abs(delta) > 1
keep_pair([2, 1, 2, 1])  # → (1.5, True)
keep_pair([1, 0, 1, 0])  # → (0.5, False)
  • Helpful model: 66,000 preference pairs of non-reasoning questions, with chosen and rejected responses of comparable length across the dataset to avoid a length bias. Same architecture as DeepSeek-R1 plus a head that outputs a score. Trained for one epoch, batch size 256, learning rate 6 × 10−6.
  • Safety model: 106,000 prompts with model-written responses labelled safe or unsafe, trained point by point rather than in pairs.
  • Each general question belongs to either the safety set or the helpfulness set, and gets that set's reward.

Why it matters

This is the classic RLHF reward, with its classic weakness, which the paper met (3.2 and Supplementary B.5). The random order and the length balancing are two standard defences against a judge's known biases; the LLM-as-a-judge companion measures those biases.

3.2 Training details · original

Everyday picture

Two rounds of coaching with different goals. The first is drills: maths, code and logic, with a rule to stay in one language. The second mixes drills with conversation practice judged by the trained judges, but only briefly, because a student coached too long by a judge learns to please the judge.

The first RL stage, and the language reward

Same settings as R1-Zero (learning rate 3 × 10−6, KL coefficient 0.001, temperature 1, 16 samples of up to 32,768 tokens, 512 outputs per step, reference reset every 400 steps). The paper gives the GRPO clip ratio for this stage as ε = 10, and notes that the clip ratio matters a great deal: too low and the gradients of many tokens are cut off, too high and training becomes unstable. Read literally, a range of 1 ± 10 would almost never clip a ratio, which is hard to square with that remark; the paper gives no other value, so treat this one with care. To curb language mixing, a new reward is added:

In words: “the language reward is the share of words in the chain of thought that are in the target language.”

With the numbers (illustrative): an English question whose 40-word chain of thought contains 34 English words scores 34 / 40 = 0.85; a chain entirely in English scores 1.

In Python:

# each word of a 40-word chain: True if it is in the target language (illustrative)
in_target = [True] * 34 + [False] * 6
# Num(Words_target) / Num(Words)
sum(in_target) / len(in_target)  # → 0.85

Supplementary B.6 shows the trade: without it, language consistency decays as training goes on; with it, consistency holds, maths scores stay level and coding scores dip slightly. The paper keeps it because people find the output more readable.

The second RL stage: everything at once

In words: “a question's reward is the rule-based score if it is a reasoning question, the reward model's score plus the format score if it is a general one, plus the language score either way.”

With the numbers (illustrative): a maths answer that is right and well formatted, in one language: 2 + 0 + 1 = 3. A general reply with reward-model score 0.6, correct format and 95% target-language words: 0 + (0.6 + 1) + 0.95 = 2.55.

In Python:

def reward(kind, rule=0.0, reward_model=0.0, fmt=0.0, language=1.0):
    reasoning = rule if kind == "reasoning" else 0.0
    general = reward_model + fmt if kind == "general" else 0.0
    # Reward_reasoning + Reward_general + Reward_language
    return reasoning + general + language
reward("reasoning", rule=2)  # → 3.0
round(reward("general", reward_model=0.6, fmt=1, language=0.95), 2)  # → 2.55

This stage runs 1,700 steps at a lower temperature, 0.7, because higher temperatures here gave incoherent text. General instruction data and the preference-based rewards enter only in the last 400 steps: the paper found that training longer on the model-based reward invites reward hacking.

Why it matters

The mixture is the engineering lesson: keep the uncheatable rule-based rewards on for the whole run, and let a learnable, hackable judge in only briefly at the end. The lesson's section on reward hacking shows why the second rule exists: optimise a learned reward long enough and true quality turns down while the reward keeps rising.

4 Experiment · original

Everyday picture

Checking a student after each term: did the neat-handwriting lessons cost them any maths? Did conversation practice help their essays? The paper scores every checkpoint in its pipeline on the same tests.

Tiny example

On IF-Eval, a test of following formatting instructions, R1-Zero scores 46.6; after the cold start, Dev1 scores 71.7. But on AIME, Dev1 drops from R1-Zero's 77.9 to 59.0: a few thousand readable examples taught manners and cost reasoning, which the next RL stage (Dev2, 74.0) buys back.

Benchmark:

Reading it: each bar is one checkpoint of the pipeline, from R1-Zero at the top to the final R1 at the bottom, on one benchmark, axis 0 to 100. Start with AIME: the dip at Dev1 and the recovery through Dev2 and Dev3 to 79.8. Switch to IF-Eval and AlpacaEval 2.0: the general-purpose scores climb at almost every stage, and AlpacaEval jumps most in the final RL stage (62.1 to 87.6), when preference rewards enter. Switch to GPQA Diamond: R1-Zero is the best checkpoint (75.8), and the final R1 scores 71.5; the paper's text does not comment on this dip. Numbers from Table 3 of DeepSeek-AI (2025), selected rows.

Why it matters

Each stage buys something and costs something, and the table makes the costs visible. The paper's reading: reasoning-oriented RL lifts reasoning a lot and preference benchmarks only a little; general data and preference rewards lift the latter; the final stage improves maths and code only marginally because most of that work was already done.

5 Ethics and safety statement · original

Everyday picture

A student who has learned to plan carefully plans everything carefully, including things they shouldn't. Better reasoning makes a harmful answer more workable, and an openly released model can be fine-tuned to strip its safety training out.

What the paper reports

  • Measured alone, R1's safety is “moderate”, comparable to GPT-4o (May 2024); combined with the external risk control system DeepSeek deploys, the paper calls it “superior”.
  • On six public safety benchmarks (Supplementary D.3, Table 9), R1 averages 95.0 with the risk control system and 85.9 without it. Its weakest benchmark is HarmBench (89.3 with the risk control system, 35.0 without); the paper traces the weakness to intellectual-property requests, such as song lyrics, which R1 does not refuse.
  • Across 50 languages, R1's safety score is 85.9% with risk control and 74.2% without.
  • Under jailbreak attacks every model tested got less safe; the two reasoning models, R1 and o1, leaned most heavily on refusals from their risk control systems (rejection rates of 79.8% and 87.3%).

Why it matters

The paper's advice to anyone serving the open weights is to add a similar risk control system. The broader point stands for any open model: safety measured on the raw model and safety measured on a deployed system are different numbers, and a report should say which it gives. The guardrails lesson builds checks of this kind.

6 Conclusion, limitation, and future work · original

“We believe that the key to unlocking this potential lies not in large-scale human annotation but in the provision of hard reasoning questions, a reliable verifier, and sufficient computational resources for reinforcement learning.”DeepSeek-AI (2025), §6

Everyday picture

A powerful engine with a narrow road. It goes very fast wherever there is a clear finish line, and has no way forward where nobody can say what finishing means.

The limitations the paper lists

  • Structured output and tools. Weaker at structured output than existing models; cannot use search engines or calculators yet.
  • Token efficiency. It spends fewer tokens on easy questions and more on hard ones, but still overthinks simple questions at times.
  • Language mixing. Optimised for Chinese and English; asked in another language, it may reason and answer in English.
  • Prompt sensitivity. Few-shot prompts consistently make it worse; the paper recommends describing the problem directly, with no examples.
  • Software engineering. Slow test runs kept large-scale RL off these tasks, so R1 improved little over DeepSeek-V3 on them.
  • Reward hacking. Pure RL needs reliable rewards; for writing and other tasks without a rule-based check, a model-based reward gets exploited as training goes on. For such tasks R1 uses human-annotated SFT data and only hundreds of RL steps.

Why it matters

The limits map the frontier the paper opened: RL is strong wherever a verifier exists, and the open problem is rewards for tasks that can't be checked. The paper also points to tool use during reasoning, from compilers and search engines to real-world experiments, as a way to widen what can be checked.

Supplementary material, selected · original

The second version adds long supplementary sections, A to J. This page walks through the ones that explain the method or change how its results should be read; each links to the original.

A: background, and GRPO against PPO · original

Everyday picture

PPO grades each attempt against a forecast of how well it should have gone; GRPO grades it against the other attempts. A forecast needs a forecaster, trained and fed like any other model.

What Supplementary A says

  • The base model. DeepSeek-V3-Base is a mixture-of-experts model with 671 billion parameters, 37 billion active per token, pretrained on 14.8 trillion tokens of web pages and e-books. The paper notes that some web pages contain answers written by OpenAI models, so the base model may have absorbed knowledge from them indirectly, and that its large volume of maths and code gives RL plausible solutions to select among.
  • Why skip SFT first. Human-written solutions often omit reflection and verification, so imitating them first may limit what RL can find (A.2).
  • The value model's problem. PPO usually computes advantages with generalized advantage estimation and a value model of similar size to the policy. That model must predict the final reward from a partial answer, which is hard when only the outcome is rewarded, and harder still when a long chain may later revise what it wrote.
  • The KL placement. PPO adds a per-token KL penalty to the reward, so the penalty accumulates with length and may discourage longer answers; GRPO adds its estimator directly to the loss.
  • The comparison. On the MATH task with DeepSeek-Coder-V2-Lite (16B mixture-of-experts, 2.4B active), PPO with the common default GAE setting λ = 0.95 did considerably worse than GRPO; tuned to λ = 1.0, it came close. GRPO reached that level without the tuning or the second model.

Why it matters

The fair summary is not “GRPO beats PPO” but “GRPO matches a carefully tuned PPO with less machinery”. The lesson's comparison diagram of PPO and GRPO draws the same trade.

B: data, hyper-parameters, cost and reward hacking · original

Everyday picture

The recipe card on the back of the box: the ingredients in grams, the oven settings, the price, and a warning about what burns.

The data

  • RL prompts (Table 4): 26K maths, 17K code, 22K STEM multiple choice, 15K logic, and 66K general. Maths answers are numbers, expressions or equations checked against a reference (proofs are excluded as too hard to check), rewarded 1 or 0.
  • Cold start: R1-Zero samples at temperature 1.0, filtered for right answers and readable format (SymPy compares maths answers; rules catch repetition and language mixing), rewritten into a first-person conversational style by DeepSeek-V3 and checked by people. Code test cases are generated by a model and filtered against known right and wrong submissions.
  • The 800K SFT set: about 600K reasoning samples kept by rejection sampling (some judged by DeepSeek-V3 against a reference answer) and about 200K non-reasoning samples, 804,745 in total, averaging 5,355 tokens each. Most are single-turn, which the paper notes may limit multi-turn conversation.
  • SFT settings: 2 to 3 epochs on DeepSeek-V3-Base, learning rate decaying from 5 × 10−5 to 5 × 10−6, context 32,768 tokens, batch 128.

Tiny example: what it cost

R1-Zero trained on 64 × 8 = 512 H800 GPUs for about 198 hours: 512 × 198 = 101,376 GPU hours, which the paper rounds to 101K. R1 used the same GPUs for about 80 hours (41K GPU hours), and making the SFT data took 5K. At the paper's assumed rental price of 2 US dollars per GPU hour, that is about $294K in all.

In words: “the compute used is the number of GPUs times the hours they ran; the rental cost is that times the price per GPU hour.”

With the numbers: 512 × 198 = 101,376 (R1-Zero); 512 × 80 = 40,960 (R1); with 5,000 for data, about 147K GPU hours, and 147K hours at 2 dollars each come to $294K.

In Python:

n_gpu, price = 64 * 8, 2
# R1-Zero: about 198 hours; R1: about 80 hours
zero, r1 = n_gpu * 198, n_gpu * 80
zero, r1  # → (101376, 40960)
# plus about 5K GPU hours to create the SFT data, then the paper rounds each part to thousands
total_k = round(zero / 1000) + 5 + round(r1 / 1000)
total_k, total_k * price  # → (147, 294)

One inconsistency: the text describes R1's run as “about 4 days, or roughly 80 hours”, but 4 days is 96 hours. The table's 41K GPU hours matches 80 hours on 512 GPUs, so 80 hours is the figure used here. This is also only the final runs, not the research that preceded them: the paper says smaller 30B-parameter experiments on A100 GPUs came first.

Reward hacking, observed

Training with the helpfulness reward model, the reward kept rising over 700 steps while the Codeforces pass rate, measured separately, fell (Figure 6). That is the reason general data and preference rewards enter only in the last 400 of the second stage's 1,700 steps.

Why it matters

The cost figures made headlines, but read carefully they cover the reasoning training only, on top of a base model whose own pretraining cost is reported elsewhere. The reward-hacking curve is the textbook picture of a proxy reward coming apart from the goal; the lesson's optimise_lengths reproduces it in miniature, and reasoning_cost prices the thinking itself.

C: the self-evolution of R1-Zero · original

Everyday picture

A diary of the workbook student: which problems they could do in which week, and how often they wrote “hmm, let me check”.

What was measured

  • By difficulty (MATH, levels 1 to 5). Easy levels 1 to 3 reach 0.90 to 0.95 accuracy early and stay there; level 4 climbs from about 0.78 to 0.95, and level 5 from about 0.55 to 0.90. Level 1 has only 43 of the 500 problems, so its accuracy of 95 to 97% means just one or two misses, mostly geometry.
  • Reflective words. Three people chose ten words (“wait”, “mistake”, “however”, “but”, “retry”, “error”, “verify”, “wrong”, “evaluate”, “check”); their count rises 5 to 7 times over training.
  • “Wait”, specifically. Nearly absent early, used occasionally between steps 4,000 and 7,000, then spiking after step 8,000. The paper reads this as different kinds of reflection appearing at particular stages.

Why it matters

This is the evidence behind “emergent reasoning”, and it is worth reading as what it is: counts of words, not of verified reflections. The counts are striking, and the accuracy gains on the hardest problems are the stronger evidence that the longer chains do real work.

D: how the results were measured · original

Everyday picture

One attempt at a test is a noisy measure of a student. Averaging over many attempts is fairer, and asking “would the student's most common answer be right?” measures something else again.

Tiny example

Four sampled answers to one AIME problem: right, wrong, right, right. Their pass@1 is 3/4 = 0.75. Their majority answer is right, so cons@k is 1.

In words: “sample k answers to a question and report the share that are right: the chance that a single sampled answer is right.”

With the numbers: k = 4, correctness (1, 0, 1, 1), pass@1 = 3/4 = 0.75.

In Python:

# p_i: 1 if the i-th sampled answer is right
p = [1, 0, 1, 1]
k = len(p)
# (1/k) Σ_i p_i
sum(p) / k  # → 0.75
  • Samples use temperature 0.6 and top-p 0.95, because greedy decoding of long reasoning outputs repeated itself more and varied from checkpoint to checkpoint. Outputs are capped at 32,768 tokens.
  • Decontamination: training text containing any 10-word sequence from a test question or solution was removed (about six million texts in maths alone), and maths RL prompts came only from competitions before 2023. The paper admits this cannot catch paraphrases, so tests released before 2024 may still be contaminated.
  • Headline comparison (Table 8, selected): AIME 2024 pass@1, R1 79.8 against OpenAI-o1-1217's 79.2; MATH-500, 97.3 against 96.4; Codeforces, better than 96.3% of human competitors against o1's 96.6%; GPQA Diamond, 71.5 against 75.7.
  • People's preferences: on the crowd-voted Chatbot Arena, with a control for response style, R1 shared first place with OpenAI-o1 and Gemini-Exp-1206 on 24 January 2025, a week after release.

Why it matters

“pass@1” here is an average over many samples, not one attempt, so it is less noisy than the name suggests, and different from the pass@1 of code benchmarks that submit one program. The lesson's pass_at_n and majority_vote build the neighbouring measures.

E: more analysis, and how long it thinks · original

Everyday picture

A good student glances at an easy question and writes the answer; on a hard one, they fill pages. Spending more effort on harder problems is itself a skill.

What Supplementary E reports

  • Thinking scales with difficulty. On 366 problems from 93 maths competitions held in 2024, R1 solves 61.8% (pass@1), using 8,793 thinking tokens per problem on average: fewer than 7,000 on the easiest and more than 18,000 on the hardest. For a question like 1 + 1, under 100 tokens.
  • Against a non-reasoning model. GPT-4o solves 24.7% of the same set with 711 output tokens on average. Majority voting over 16 samples barely helps it; on AIME 2024, voting over 64 samples lifts it only from 9.3% to 13.4%. Independent samples can't build on each other or back up from a mistake.
  • R1 still benefits from sampling. Its AIME 2024 pass@64 is 90.0% against pass@1 79.8%, and majority voting lifts it to 86.7%, so the paper suggests voting or tree search can complement long reasoning.
  • Fresh tests. On AIME 2025, released after training, R1 solves 75%, against o1's 80%; its AMC 12 2024 and AIME 2025 scores together would qualify for the USA Mathematical Olympiad. It is strongest in number theory and algebra and weakest in geometry and combinatorics.
  • Each stage on hard problems. On LiveCodeBench's hard problems, accuracy rises at every stage, from 17.1% (R1-Zero) to 34.4% (R1); easy problems are solved at every stage.

One detail to read with care: the main text reports that R1-Zero reaches 86.7% on AIME 2024 with self-consistency (its Figure 1 labels this cons@16), and this section reports that majority voting lifts R1 from 79.8% to 86.7% too. Two different models reaching the same vote score is possible, but the paper does not remark on it, and Supplementary D describes AIME consensus as cons@64.

Why it matters

This is the argument for training reasoning in rather than bolting it on: the same token budget spent as one long chain that can check and revise beat many short independent attempts, by a wide margin. The self-consistency companion shows when voting helps, and the reasoning lesson's toy reproduces thinking that scales with difficulty.

F: distillation into small models · original

Everyday picture

A student taught from a master's worked solutions learns faster than one left to rediscover everything with an answer key, especially if the student is young.

Tiny example

Fine-tune Qwen2.5-32B on R1's 800K samples and it scores 72.6 on AIME 2024. Train the same 32B base with R1-Zero-style RL for over 10,000 steps instead, and it scores 47.0. Imitating the big model's reasoning beat discovering it from scratch by more than 25 points.

Reading it: AIME 2024 pass@1, axis 0 to 100; the plain bars are models distilled from R1 by fine-tuning alone (no RL), the striped bars are comparisons. Read top to bottom: even the smallest distilled model, 1.5 billion parameters, beats GPT-4o and Claude 3.5 Sonnet on this test, and accuracy climbs with size. Then compare the two 32B bars: distillation (72.6) against RL alone on the same base (47.0, level with QwQ-32B-Preview's 50.0). Numbers from Tables 8, 15 and 16 of DeepSeek-AI (2025), selected entries.

The paper draws two conclusions: distilling a strong model into a small one is cheap and works very well, while small models trained with large-scale RL alone need enormous compute and may still fall short; but going beyond what current models can do may still need stronger base models and larger-scale RL. The distilled models were trained with SFT only, 2 to 3 epochs, leaving RL on top of them to others. A separate check on Qwen2-Math-7B, released before the first reasoning models so it could not have seen their outputs, found RL alone still lifted it well above its instruction-tuned version.

Why it matters

This is why small open reasoning models appeared so quickly: the expensive discovery happens once, in a big model, and its outputs become cheap training data. The lesson on training stages covers distillation, and distillation_loss builds the classic version from a teacher's probabilities.

G: key findings, and what didn't work · original

Everyday picture

The lab notebook's honest pages: the things that were tried, looked promising, and were abandoned.

Key findings

  • The base model matters. RL from a 7B dense model and a 16B mixture-of-experts model failed to improve AIME scores: as responses grew, they repeated themselves instead of using the length. A 32B dense, a 230B and a 671B mixture-of-experts model all improved.
  • Verifiers matter. Rule-based checks, and a language model comparing a short answer with a reference, resisted hacking; neither generalises to open-ended writing.
  • Both RL and SFT are needed: RL to discover long reasoning, SFT where no reliable reward exists.

Unsuccessful attempts

  • Process reward models, which score every step: hard to define a step, hard to label its correctness, and once a learned model scores steps, it gets hacked. Useful for reranking outputs, not worth their cost inside large-scale RL.
  • Monte Carlo tree search, inspired by AlphaGo: the space of possible token sequences is vastly larger than a board game's, limiting the search makes it get stuck, and the value model that guides it is hard to train. It helps at inference with a pretrained value model; iteratively improving the model through self-search did not work.

The paper stresses that these failures do not mean the approaches cannot work, only that they did not in its experiments.

Why it matters

Negative results are rarely published, and these two shaped the field: much of the effort went to outcome-rewarded RL with simple verifiers rather than step-level reward models and search. The reasoning lesson compares outcome and process verifiers on a toy.

What it changed

Before the paperIn the paperLearn it
Reasoning taught by imitating written solutionsReasoning found by RL with a verifier, then written down and distilledtrain_reasoner
PPO with a value model and a KL penalty in the rewardGRPO: group baselines, KL in the loss, no value modeltrain_grpo
Learned reward models for everythingRule-based rewards for reasoning; learned judges only briefly, at the endverify
Test-time compute as many independent samplesTest-time compute as one long chain that checks and revises itselfself-consistency companion
The strongest reasoning models closedOpen weights for R1, R1-Zero and six distilled modelsreasoning lesson
The policy gradient behind it allReward minus a baseline, times the gradient of the answer's log-probabilityREINFORCE companion

Glossary

Every term with hover guidance on this page, in one place.