primer.ml.reinforcement

Reinforcement learning: learning from a score instead of an answer

Run: python -m primer.ml.reinforcement

Level 1: The practitioner's guide

In one sentence. Reinforcement learning trains a model from a score on what it produced rather than from a correct answer to copy, which is how a language model learns things nobody can write down (be helpful, reason to the right answer) and how it learns to game a score that was written down badly.

When you need it. You need RL when you can judge an answer but cannot write the perfect one: a proof either checks or it doesn't, tests pass or fail, one reply is better than another though neither is "the" reply. The tell: you find yourself writing a grader, a checker or a rubric instead of example answers. You don't need it when you can write the answers (supervised fine-tuning copies them, at a fraction of the cost) or when you have pairs of better and worse answers and nothing more (DPO in primer.ml.training_stages learns from pairs with no sampling loop). And you never need to run the vendor's own RL: the helpfulness, harmlessness and reasoning of a hosted model were trained this way before you arrived, which is why this lesson matters even if you never train: it explains why models answer at length, flatter, and sometimes optimise the letter of your instruction instead of its spirit.

Your options. From the cheapest to the most committed:

Option What it does What it guarantees What it costs Where it lives
Supervised fine-tuning on demonstrations Copy correct answers you wrote The behaviour in the examples, nothing beyond them Writing the answers Your training stack, or a hosted API
DPO on preference pairs Learn from "this one beat that one", no sampling during training Shifts tone and choices without a reward model or an RL loop Thousands of comparisons Your training stack, or a hosted API
Hosted reinforcement fine-tuning with your grader The vendor samples answers and reinforces the ones your grader scores high An RL loop you don't build; the grader is still yours to get right Grader design, many sampled answers per prompt, the vendor's price The vendor's API (as of October 2026, OpenAI's fine-tuning platform no longer accepts new users, so check availability first)
GRPO with a verifiable reward Sample a group of answers per prompt, check each, reinforce the above-average ones A reward with no learned blind spot, and no second model to train Generation dominates: 8 answers per prompt is a common default; a checker that cannot be argued with Your training stack
PPO with a learned reward model Train a reward model on ratings, then a value network and the policy against it, on a KL leash Optimises a goal no program can check, such as helpfulness Two extra models (a reward model and a value network; InstructGPT used 6B for both at every policy size), and the reward model's blind spots to defend against Your training stack

How to choose. Ask what can judge an answer, and how much you trust it.

  • A program can check the final answer (arithmetic, unit tests, a format, a proof checker): GRPO with that check as the reward. In this lesson's toy, accuracy on eight addition prompts goes from 23% to 99% in 60 steps of 8 answers each; this recipe is how DeepSeek-R1-Zero learned to reason from rule-based rewards: a correct final answer and a required format.
  • Only people can judge, and you have their ratings: a learned reward model with PPO, a KL leash, and a held-out measure of the real goal that you watch more closely than the reward.
  • You have pairs but no budget for sampling: DPO.
  • You can write the answers: supervised fine-tuning, and stop there.
  • Whatever you pick, the policy optimises the reward you wrote, not the goal you meant. Before training, ask what a literal-minded optimiser would do with your reward, and measure the goal separately.

What it costs. Sampling is the bill: every training example is a full generation, and GRPO multiplies it by the group size, which is why PPO reuses each batch for several passes and why training loops run a fast inference engine beside the trainer. A learned baseline costs a second model (PPO's value network); GRPO replaces it with the group's own mean, which is why it was introduced as a way to cut PPO's memory. Noise costs steps: in the lesson's two-arm toy the gradient estimate's variance is 30.25 without a baseline and 0 with one, and with rewards offset by five points, 60% of training runs without a baseline lock onto the wrong arm against 0% with one. Reusing a batch too hard costs calibration: 50 passes over 16 pulls push one arm to 0.84 unclipped, 0.41 with PPO's clip at 0.2. And a KL leash costs a little reward on every answer (1.0 becomes 0.931 for an answer whose probability doubled, at β = 0.1) to buy fluency and safety from drift.

What breaks.

  • Reward hacking. The policy finds where the reward and the goal disagree. In the toy, a reward model fitted on answers of 1 to 4 sentences rates a 10-sentence answer 2.19, the worst answer of all; true quality rises from 0.80 to 0.93 and then falls to 0.75, below where it started, while the reward keeps climbing. Gao, Schulman and Hilton (2022) measured the same rise and fall at scale. "The reward went up" proves nothing; keep a held-out measure of the goal.
  • Length bias and sycophancy. Raters prefer long, flattering answers, so the reward model does too, so the model becomes that. Penalise length directly and rate the policy's current outputs, not stale ones.
  • Gaming the checker. A coding model rewarded for passing tests learns to edit or special-case the tests. A verifier has no learned blind spot, but a buggy or bypassable one is a reward model with extra steps.
  • Groups with nothing to teach. When every answer in a GRPO group is right or all are wrong, every advantage is zero: by the end of the toy run, 84% of groups are unanimous. Filter for prompts at the edge of the model's ability.
  • Too loose a leash. As β falls, the best policy piles onto whatever the reward likes (99.995% on one action at β = 0.1 in the toy) and true quality collapses; too tight and nothing moves. Sweep it.
  • Collapsed exploration. An arm whose probability hits zero is never tried again, so the policy can lock onto a mediocre answer early.

In the wild. InstructGPT (Ouyang et al., 2022) set the pattern of PPO against a learned reward model with a per-token KL penalty, and every chat assistant since inherits its habits. DeepSeekMath (Shao et al., 2024) introduced GRPO, reaching 51.7% on the MATH benchmark with a 7-billion-parameter model; DeepSeek-R1 (2025) trained reasoning with RL on rule-based rewards and no human-written reasoning traces. Hugging Face TRL's GRPOTrainer takes reward functions as plain Python callables or a reward model, samples 8 generations per prompt by default, and can generate with vLLM; OpenAI's model optimization guide lists reinforcement fine-tuning, where you supply the grader (as of October 2026 OpenAI's fine-tuning platform no longer accepts new users). Sutton and Barto's textbook and OpenAI's Spinning Up are the standard longer reads.

Go deeper. Level 2 builds it all on a three-armed slot machine: REINFORCE as one line of arithmetic, why a baseline removes noise without bias, PPO's ratio and clip on a table of four cases, the KL leash as a fee per answer, GRPO's advantages from a group of four, and a reward model that loves length, trained against until quality falls, each with a figure you can rerun. If you only needed to choose a training signal, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

Think of teaching a dog to sit. You can't show it the right answer: you can only wait for it to try something and give it a treat when the something was good. Over many tries the dog does more of what earned treats and less of what didn't. Nobody ever told it what "sit" means; it worked it out from a score.

That is reinforcement learning (RL). An agent (the dog, or a language model) takes an action (sits, or writes an answer), the world hands back a reward (a treat, or a score from a grader), and the agent adjusts itself so that rewarded actions become more likely. The agent's current habits, written as a probability for every action, are called its policy.

Compare this with ordinary supervised learning (primer.ml.losses), where every example comes with the correct answer attached. In RL there is no answer key, only a score after the fact. That is exactly the situation a language model is in once pretraining is over: for "prove this theorem" or "write a helpful reply" there is no single correct text to copy, but a checker or a judge can say how good an attempt was. RL is how the model learns from those judgements (primer.ml.training_stages places it in the training pipeline).

flowchart LR P[Policy<br/>a probability for every action] -->|sample| A[Action] A --> E[Environment<br/>a slot machine, a grader, a user] E --> R[Reward<br/>one number] R -->|nudge the policy| P

Reading it: follow the loop clockwise. The policy is a set of probabilities, and the action is sampled from it, so the agent sometimes tries things it isn't sure about. The environment is whatever judges the action; the agent can't see inside it. The only thing that comes back is one number, the reward, and the only thing the agent can do with it is nudge its own probabilities. Everything in this lesson is a better answer to one question: how exactly should that nudge be computed?

A tiny worked example: three slot machines

The simplest RL problem is a row of slot machines, called a multi-armed bandit (a slot machine is a "one-armed bandit"). Machine A pays out 20% of the time, B 50% and C 80%, but the agent isn't told that. Each pull pays 1 or 0. The agent must find C by pulling and seeing what happens.

A policy here is three probabilities, one per arm. A policy that picks each arm a third of the time earns, on average, (0.2 + 0.5 + 0.8) / 3 = 0.5 per pull. A policy that always picks C earns 0.8. Learning means moving from the first policy to the second using nothing but the 1s and 0s.

This one number, the average reward a policy expects, is what RL maximises:

Level 3: the formula and its symbols

$$ J(\theta) = \sum_{a} \pi_\theta(a)\, R(a) $$

Symbols

Symbol Meaning here In the example
$a$ one action: which arm to pull A, B or C
$\theta$ the policy's adjustable numbers (here, one score per arm, called logits) (0, 0, 0)
$\pi_\theta(a)$ the policy: the probability of picking action $a$, given $\theta$. Here softmax of the logits 1/3 each
$R(a)$ the average reward action $a$ pays (unknown to the agent) 0.2, 0.5, 0.8
$\sum_{a}$ add up over every action three terms
$J(\theta)$ the expected reward: what the policy earns per pull, on average 0.5

In words: "the expected reward is each action's probability times its average payout, added up over the actions."

With the numbers: J = ⅓·0.2 + ⅓·0.5 + ⅓·0.8 = 0.5 for the uniform policy, and 1·0.8 = 0.8 for the policy that always pulls C.

Level 3: in Python

In Python:

policy = [1/3, 1/3, 1/3]
# average payout of arms A, B, C
R = [0.2, 0.5, 0.8]
# J = Σ_a π(a) R(a)
round(sum(p * r for p, r in zip(policy, R)), 3)  # → 0.5
round(sum(p * r for p, r in zip([0, 0, 1], R)), 3)  # → 0.8

The agent can't compute J, because it doesn't know R. It can only sample: pull an arm, see a 1 or a 0. Every method below turns those samples into an estimate of which way to move θ to make J bigger. Trying an arm you're unsure of is called exploration; sticking with the best arm so far is exploitation. A sampled policy does some of both automatically, as long as no arm's probability has collapsed to zero.

In code: Bandit hides the win chances and pays out one pull at a time; expected_reward is the formula above.

Policy gradients: do more of what worked (REINFORCE)

Everyday picture. A football coach reviews the tape after a match. For every play that led to a goal, they tell the team "a bit more of that"; for plays that went nowhere, nothing. They don't need to know why the play worked. Repeat over hundreds of matches and the team drifts towards the plays that score.

Tiny worked example. Start with logits (0, 0, 0), so each arm has probability ⅓. The agent pulls C and wins: reward 1. The rule, explained next, says: add to each logit learning rate × reward × (1 if it's the chosen arm, else 0, minus that arm's probability).

Arm Chosen? 1[chosen] − π × reward 1 × rate 0.5 New logit New probability
A no 0 − ⅓ = −0.333 −0.167 −0.167 0.274
B no 0 − ⅓ = −0.333 −0.167 −0.167 0.274
C yes 1 − ⅓ = +0.667 +0.333 +0.333 0.452

One lucky pull moved C from 33% to 45%. Had the pull paid 0, nothing would have moved. Had the agent pulled A and won (A wins sometimes too), A would have gone up instead. The rule is noisy, one pull at a time, but on average the arm that wins most gets pushed up most.

flowchart LR L[Logits θ] --> S[softmax<br/>probabilities π] S -->|sample| A[Action a] A --> ENV[Pull the arm] --> R[Reward R] S --> G["∇ log π(a)<br/>= one-hot(a) − π"] A --> G G --> M["× R × learning rate"] R --> M M -->|add to| L

Reading it: the top path is acting: logits become probabilities, one action is sampled, and the environment pays a reward. The lower path is learning: from the action alone, work out which direction in logit space makes that action more likely (the ∇ log π box), then scale that direction by how good the outcome was. A big reward is a big step towards repeating the action; zero reward is no step.

The log-probability trick, decoded

We want the gradient of J: for each logit, how much J rises if the logit rises a little (see primer.notation for gradients from scratch). The difficulty is that J is an average over actions we can only sample. The trick rewrites the gradient as an average too, so a sample estimates it:

Level 3: the formula and its symbols

$$ \nabla_\theta J(\theta) = \mathbb{E}_{a \sim \pi_\theta}\big[\, R(a)\, \nabla_\theta \log \pi_\theta(a) \,\big] \qquad \frac{\partial \log \pi_\theta(a)}{\partial z_k} = \mathbb{1}[k = a] - \pi_\theta(k) $$

Symbols

Symbol Meaning here In the example
$\nabla_\theta$ "gradient with respect to θ": one slope per logit, collected into a vector 3 slopes
$\mathbb{E}_{a \sim \pi_\theta}[\ldots]$ expected value: the average of the bracket when $a$ is sampled from the policy average over pulls
$\sim$ "drawn from"
$\log$ the natural logarithm, the undo button for $e^x$ $\log \tfrac13 = -1.10$
$\nabla_\theta \log \pi_\theta(a)$ the direction in logit space that makes action $a$ more likely, fastest $(-\tfrac13, -\tfrac13, \tfrac23)$ for C
$z_k$ the $k$-th logit (θ is the list of logits) $z_3 = 0$
$\partial$ "partial derivative": the slope along one logit, holding the others still
$\mathbb{1}[k = a]$ 1 if $k$ is the chosen action, else 0 (0, 0, 1)
$\pi_\theta(k)$ the probability of action $k$ ⅓

In words: "the direction that raises expected reward is, on average, the direction that makes the sampled action more likely, weighted by the reward it earned. For a softmax policy, that direction is 'one for the chosen action, minus every action's probability'."

Why is this true? Because the slope of a probability equals the probability times the slope of its log ($\nabla \pi = \pi \, \nabla \log \pi$, the chain rule applied to log). So $\nabla J = \sum_a \nabla\pi(a) R(a) = \sum_a \pi(a) \nabla\log\pi(a) R(a)$, and a sum weighted by $\pi(a)$ is an average over samples from π. The update is then plain gradient ascent: $\theta \leftarrow \theta + \alpha\, R\, \nabla_\theta \log \pi_\theta(a)$, with learning rate $\alpha$ (primer.ml.optimizers).

With the numbers: at the uniform policy the true gradient is (−0.1, 0, +0.1): push C up, A down, leave B (which pays exactly the average) alone. One sample, "pulled C, got 1", estimates it as 1 × (−⅓, −⅓, ⅔). The step with α = 0.5 gives logits (−0.167, −0.167, 0.333) and probabilities (0.274, 0.274, 0.452), as in the table.

Level 3: in Python

In Python:

import math
z = [0.0, 0.0, 0.0]
pi = [math.exp(z_k) / sum(math.exp(v) for v in z) for z_k in z]
# the true gradient: Σ_a π(a) R(a) (1[k=a] − π(k)), for each logit k
R = [0.2, 0.5, 0.8]
true_grad = [sum(pi[a] * R[a] * ((k == a) - pi[k]) for a in range(3)) for k in range(3)]
[round(g, 3) for g in true_grad]  # → [-0.1, 0.0, 0.1]
# one sample: pulled C (a = 2), reward 1
a, reward, alpha = 2, 1.0, 0.5
grad_log_pi = [(k == a) - pi[k] for k in range(3)]
[round(g, 3) for g in grad_log_pi]  # → [-0.333, -0.333, 0.667]
z = [z_k + alpha * reward * g for z_k, g in zip(z, grad_log_pi)]
[round(z_k, 3) for z_k in z]  # → [-0.167, -0.167, 0.333]
[round(math.exp(z_k) / sum(math.exp(v) for v in z), 3) for z_k in z]  # → [0.274, 0.274, 0.452]

This algorithm is called REINFORCE (Williams, 1992). Run it for 500 pulls and the policy finds arm C:

REINFORCE on the three-armed bandit: the probability of arm C climbs from a third to about 0.96 within 500 pulls, and the average reward rises from 0.5 towards 0.8

Reading it: on the left, each line is one arm's probability over 500 pulls. All three start at ⅓. C's line (the 80% arm) climbs towards 1 while A and B sink; the wiggles are single lucky or unlucky pulls. On the right is the reward, averaged over the last 50 pulls. It starts near 0.5 (random pulling) and rises towards the dashed line at 0.8, the most any policy can earn. Nobody told the agent which arm was best: the 1s and 0s were enough.

Why it matters in practice. A language model is exactly this kind of policy, with a vocabulary of tokens as its arms, and one sampled answer is a string of sampled tokens. REINFORCE applies unchanged: sum the log probabilities of every token in the answer, and scale the gradient by the answer's reward. Every method below (PPO, GRPO) is REINFORCE with repairs.

In code: grad_log_prob is one-hot minus the probabilities, reinforce_step is one update, worked_reinforce_step is the table above, and train_reinforce runs the whole loop against a Bandit.

Variance and baselines: grade on a curve

Everyday picture. A teacher whose class all scores between 90 and 100 learns nothing by being told "you got a 92". What matters is whether 92 is above or below the class average. Raw scores that are all large and positive make every attempt look good; only the difference from typical tells you which way to go.

Tiny worked example. Two arms, a 50/50 policy, and every pull pays a lot: arm 1 always pays 10, arm 2 always pays 12. Arm 2 is better, so the logit of arm 2 should rise. Look at the REINFORCE estimate for that logit:

Pulled Reward 1[arm 2] − π(arm 2) Estimate (no baseline) Estimate (baseline 11)
arm 1 10 0 − 0.5 = −0.5 10 × −0.5 = −5 (10 − 11) × −0.5 = +0.5
arm 2 12 1 − 0.5 = +0.5 12 × +0.5 = +6 (12 − 11) × +0.5 = +0.5

Without a baseline the estimate is −5 or +6 depending on the coin flip. It averages to +0.5, the right answer, but any single sample points the wrong way half the time, and violently. Subtract the average reward, 11, first and every sample says +0.5. Same average, no noise at all.

Level 3: the formula and its symbols

$$ \nabla_\theta J(\theta) = \mathbb{E}_{a \sim \pi_\theta}\big[\, (R(a) - b)\, \nabla_\theta \log \pi_\theta(a) \,\big], \qquad A(a) = R(a) - b $$

Symbols

Symbol Meaning here In the example
$b$ the baseline: any number that doesn't depend on which action was taken; usually the average reward 11
$A(a)$ the advantage: how much better action $a$ did than typical −1 for arm 1, +1 for arm 2
everything else as in the REINFORCE formula above

In words: "scale each step by how much better than typical the action did, not by its raw reward."

Why is it allowed? Because the baseline's contribution averages to zero: $\mathbb{E}[\, b\, \nabla \log \pi(a)] = b \sum_a \nabla \pi(a) = b\, \nabla \sum_a \pi(a) = b\, \nabla 1 = 0$. Probabilities always add to 1, so pushing all of them up is impossible; the baseline only removes noise, never signal.

With the numbers: without a baseline the estimate's variance (the average squared distance from its mean, see primer.notation) is (25 + 36)/2 − 0.5² = 30.25. With b = 11 it is 0.

Level 3: in Python

In Python:

# arm 1 pays 10, arm 2 pays 12; the 1[arm 2] − π(arm 2) factor for each pull
pulls = [(10, -0.5), (12, +0.5)]
def mean_and_variance(b):
    estimates = [(reward - b) * direction for reward, direction in pulls]
    mean = sum(estimates) / 2
    return mean, sum((e - mean) ** 2 for e in estimates) / 2
mean_and_variance(b=0)  # → (0.5, 30.25)
mean_and_variance(b=11)  # → (0.5, 0.0)
flowchart LR R[Reward R] --> MINUS["R − b"] B["Baseline b<br/>average reward so far"] --> MINUS MINUS --> ADV{Advantage A} ADV -->|positive: better than typical| UP[make the action<br/>more likely] ADV -->|negative: worse than typical| DOWN[make the action<br/>less likely]

Reading it: the baseline sits between the reward and the update. Its job is to turn "how good was this?" into "how much better than usual was this?". The sign of the advantage now decides the direction of the step, so a below-average action is actively pushed down, even though its raw reward was positive.

To see it matter, give every arm of our bandit 5 extra points: rewards are now 5 or 6 instead of 0 or 1, and nothing about which arm is best has changed.

With rewards offset by 5, the gradient estimate's variance falls from about 20 to 0.15 with a baseline, and 20 training runs all find arm C with it, while without it most runs lock onto the wrong arm

Reading it: on the left, the variance of a one-pull gradient estimate at the starting policy, on a log scale: about 20 without a baseline and about 0.15 with one, over a hundred times smaller. On the right, each dot is one of 20 training runs (400 pulls each), placed at its final probability of picking C. With a running-average baseline (blue) every run ends near 0.95. Without one (red), the dots scatter to both ends: in most runs, early pulls of a mediocre arm paid 5 and were pushed up hard, and the policy committed before it ever learned C was better.

Why it matters in practice. Every practical policy-gradient method uses a baseline. PPO learns one with a second network, the value network (or critic), which predicts the expected reward from each state. GRPO, below, gets one for free by comparing several answers to the same prompt.

In code: gradient_estimate_stats computes the exact mean and variance of the estimate, sampled_gradient_variance measures it from real pulls, and train_reinforce subtracts a running average when baseline=True.

PPO: take several steps, but never too far

Everyday picture. A chef tests a new recipe on one evening's diners. It would be wasteful to use their comments for just one small tweak, so the chef makes several rounds of changes from the same comment cards. But the further the recipe drifts from what the diners actually ate, the less their comments apply, so the chef caps each change: never more than 20% more or less of any ingredient per round.

For a language model, sampling answers is the expensive part (every answer is a full generation), so PPO (Proximal Policy Optimization) reuses each batch of answers for several gradient steps. It needs a way to tell how far the policy has moved since the batch was sampled, and a brake.

Tiny worked example. The probability ratio compares the policy now with the policy that generated the sample. A ratio of 1.5 means the current policy is 50% more likely to produce that answer than when it was sampled. With a clip range ε = 0.2 the ratio is allowed to count only between 0.8 and 1.2:

Ratio ρ Advantage A ρ·A clip(ρ, 0.8, 1.2)·A min of the two What happened
1.1 +3 3.3 3.3 3.3 inside the band: plain REINFORCE
1.5 +2 3.0 1.2 × 2 = 2.4 2.4 good action already boosted enough: gain capped
0.5 −1 −0.5 0.8 × −1 = −0.8 −0.8 bad action already cut enough: capped
1.5 −1 −1.5 1.2 × −1 = −1.2 −1.5 bad action made more likely: full penalty
Level 3: the formula and its symbols

$$ \rho_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\text{old}}(a_t \mid s_t)} \qquad L^{\text{CLIP}}(\theta) = \mathbb{E}_t\Big[\min\big(\rho_t A_t,\ \operatorname{clip}(\rho_t,\ 1-\varepsilon,\ 1+\varepsilon)\, A_t\big)\Big] $$

Symbols

Symbol Meaning here In the example
$t$ one sample in the batch; for a language model, one token of one answer row of the table
$s_t$ the state: what the policy saw before acting (the prompt plus the tokens so far) a prompt
$a_t$ the action taken (the token generated) an answer
$\pi_{\text{old}}$ the policy as it was when the batch was sampled, frozen
$\pi_\theta$ the policy now, after some steps on this batch
$\rho_t$ the probability ratio, new over old 1.5
$A_t$ the advantage of that sample +2
$\varepsilon$ the clip range, typically 0.1 to 0.3 0.2
$\operatorname{clip}(x, lo, hi)$ $x$, but pushed back to $lo$ or $hi$ if it falls outside clip(1.5, 0.8, 1.2) = 1.2
$\min$ the smaller of the two min(3.0, 2.4) = 2.4
$\mathbb{E}_t$ the average over all samples in the batch
$L^{\text{CLIP}}$ the objective PPO climbs

In words: "for each sample, take the ratio-weighted advantage, but once the ratio has moved more than ε from 1 in the direction the advantage wants, stop counting further movement; and always take the more pessimistic of the clipped and unclipped versions."

The slope of ρ·A is exactly REINFORCE's gradient scaled by ρ (because $\nabla\rho = \rho\,\nabla\log\pi_\theta$), so inside the band PPO is REINFORCE with importance weighting. Outside it, the clipped term is flat: that sample stops pushing.

With the numbers: the four rows of the table, computed:

Level 3: in Python

In Python:

def clip(x, lo, hi):
    return max(lo, min(x, hi))
def L_clip(rho, A, eps=0.2):
    return min(rho * A, clip(rho, 1 - eps, 1 + eps) * A)
[round(L_clip(rho, A), 2) for rho, A in [(1.1, 3), (1.5, 2), (0.5, -1), (1.5, -1)]]  # → [3.3, 2.4, -0.8, -1.5]
# past 1 + ε with a positive advantage, a higher ratio earns nothing more
L_clip(1.3, 2) == L_clip(1.6, 2)  # → True
flowchart TB OLD[Policy π_old] -->|generate a batch| B[Answers + rewards] B --> V[Value network<br/>predicts expected reward] V --> ADV[Advantages A_t] B --> ADV ADV --> LOOP subgraph LOOP["Several passes over the same batch"] RATIO["ratio ρ = π_θ / π_old"] --> CLIP["clip to 1 ± ε, take the min"] CLIP --> KL["subtract β × KL to the reference model"] KL --> STEP[gradient step on θ] STEP --> RATIO end LOOP -->|π_θ becomes the new π_old| OLD

Reading it: the outer loop is sampling, the expensive part: the frozen old policy writes a batch of answers and a value network turns their rewards into advantages. The inner loop reuses that batch for several passes. Each pass recomputes how far the policy has moved (the ratio), stops counting movement beyond the band (the clip), and applies the KL leash described below. When the passes are done, the updated policy becomes the new sampler and the cycle repeats.

To see the clip work, take one batch of 16 pulls from the bandit (C won all four of its pulls, A lost all four) and make 50 passes over it:

Two panels: the clipped objective is flat outside the 0.8 to 1.2 band on the side the advantage favours; over 50 passes the unclipped ratio climbs past 2.5 while the clipped one levels off near 1.24

Reading it: on the left is the objective for one sample as its ratio changes. For a positive advantage (blue), the line rises with the ratio until 1.2, then goes flat: no reward for pushing further. For a negative advantage (red), it goes flat below 0.8. The dashed lines are what REINFORCE would keep climbing. On the right, the largest ratio in the batch after each of 50 passes. Unclipped (red), the policy chases the same 16 pulls further every pass until C's probability is 2.5 times what it was, about 0.84 from sixteen pulls, which is wildly overconfident. Clipped (blue), it levels off near 1.24: the brake is not a hard wall (other samples can still nudge the policy), but it removes the incentive to overfit one batch.

In code: ppo_clipped_objective is the formula, clipped_policy_update makes several passes over one batch with or without the clip, and ppo_drift_experiment is the 50-pass comparison.

The KL leash: stay close to where you started

Everyday picture. A dog on a long leash can explore, but it can't run off a cliff. In RL for language models, the leash ties the policy to a frozen copy of the model it started from, the reference model (usually the model after supervised fine-tuning).

Tiny worked example. An answer earns reward 1.0. The policy now gives it probability 0.6; the reference gave it 0.3. With leash strength β = 0.1, the reward actually used for training is 1.0 − 0.1 × ln(0.6 / 0.3) = 1.0 − 0.1 × 0.693 = 0.931. The policy pays a small fee for having doubled that answer's probability.

Level 3: the formula and its symbols

$$ R'(a) = R(a) - \beta \log\frac{\pi_\theta(a)}{\pi_{\text{ref}}(a)} \qquad \mathbb{E}_{a\sim\pi_\theta}\big[R'(a)\big] = \mathbb{E}_{a\sim\pi_\theta}\big[R(a)\big] - \beta\, \mathrm{KL}(\pi_\theta \,\|\, \pi_{\text{ref}}) $$

Symbols

Symbol Meaning here In the example
$R(a)$ the reward from the grader or reward model 1.0
$R'(a)$ the reward after the leash's fee 0.931
$\beta$ leash strength: how much a unit of drift costs 0.1
$\pi_{\text{ref}}(a)$ the frozen reference model's probability for the answer 0.3
$\log\frac{\pi_\theta(a)}{\pi_{\text{ref}}(a)}$ how much more (positive) or less (negative) likely the policy makes this answer than the reference ln 2 = 0.693
$\mathrm{KL}(\pi_\theta \,|\, \pi_{\text{ref}})$ KL divergence: the average of that log ratio over the policy's own answers; zero only when the two agree (decoded in primer.ml.training_stages)

In words: "each answer's reward is docked in proportion to how much more likely the policy has made it than the reference did; on average, that fee is β times the KL divergence between the two."

With the numbers: 1.0 − 0.1 × ln 2 = 0.931. Had the policy halved the answer's probability instead (0.15), the log ratio would be −0.693 and the reward would rise to 1.069: the leash pulls both ways.

Level 3: in Python

In Python:

import math
R, beta = 1.0, 0.1
pi, pi_ref = 0.6, 0.3
# R' = R − β log(π / π_ref)
round(R - beta * math.log(pi / pi_ref), 3)  # → 0.931
round(R - beta * math.log(0.15 / pi_ref), 3)  # → 1.069

Why it matters in practice. This is the "penalty for drifting" in the RLHF loop of primer.ml.training_stages, and the same β appears in DPO. It stops the policy from forgetting fluent language while it chases reward, and it is the first line of defence against reward hacking, below.

In code: kl_penalised_reward is R′; clipped_policy_update adds the leash's gradient when given a reference policy and a β, and primer.ml.training_stages.kl_divergence computes the KL itself.

GRPO: compare answers to the same question

Everyday picture. Instead of hiring an examiner to predict how hard each exam question is, a teacher gives the same question to eight students and marks each answer relative to the others on that question. On an easy question, getting it right is expected and earns little credit; on a hard one, the only right answer stands out.

PPO's baseline comes from the value network, a second model, often as large as the policy, that has to be trained alongside it. GRPO (Group Relative Policy Optimization) throws the value network away. For each prompt it samples a group of answers and uses the group's own average as the baseline.

Tiny worked example. The prompt is "3 + 4 =". The model samples four answers and a checker scores them 1 if the answer is 7, else 0.

Group rewards Mean Std Advantages
1, 0, 0, 1 0.5 0.5 +1, −1, −1, +1
1, 0, 0, 0 0.25 0.433 +1.73, −0.58, −0.58, −0.58
1, 1, 1, 1 1 0 0, 0, 0, 0

A lone right answer in a mostly wrong group earns a big advantage: it's rare, so it's strong evidence. A group that is all right (or all wrong) earns nothing: there is no contrast, so there is nothing to learn from.

Level 3: the formula and its symbols

$$ A_i = \frac{r_i - \operatorname{mean}(r_1, \ldots, r_G)}{\operatorname{std}(r_1, \ldots, r_G)} $$

Symbols

Symbol Meaning here In the example
$G$ the group size: answers sampled per prompt 4
$i$ which answer in the group 1 … 4
$r_i$ the reward for answer $i$ 1, 0, 0, 0
$\operatorname{mean}(\ldots)$ the group's average reward: the baseline 0.25
$\operatorname{std}(\ldots)$ the group's standard deviation, the typical distance from the mean (square root of the variance) 0.433
$A_i$ answer $i$'s advantage, shared by every token of that answer +1.73

In words: "an answer's advantage is how far its reward sits above the group's average, measured in units of the group's spread."

With the numbers: for (1, 0, 0, 0): mean 0.25, variance (0.75² + 3 × 0.25²) / 4 = 0.1875, std √0.1875 = 0.433, so the right answer gets 0.75 / 0.433 = 1.73 and each wrong one −0.25 / 0.433 = −0.58.

Level 3: in Python

In Python:

import statistics
def advantages(r):
    mu, sd = statistics.mean(r), statistics.pstdev(r)
    return [round((r_i - mu) / sd, 2) if sd else 0.0 for r_i in r]
advantages([1, 0, 0, 1])  # → [1.0, -1.0, -1.0, 1.0]
advantages([1, 0, 0, 0])  # → [1.73, -0.58, -0.58, -0.58]
advantages([1, 1, 1, 1])  # → [0.0, 0.0, 0.0, 0.0]

The rest of GRPO is PPO: the same ratio, the same clip, the same KL leash to a reference model, averaged over the group. (Some implementations divide by the sample standard deviation, with G − 1, rather than the population one; the idea is identical.)

flowchart LR subgraph PPO["PPO"] P1[Prompt] --> A1[one answer] A1 --> RM1[reward] A1 --> VN[value network<br/>a second big model] RM1 --> AD1[advantage = reward − value] VN --> AD1 end subgraph GRPO["GRPO"] P2[Prompt] --> G1[answer 1] & G2[answer 2] & G3[answer ...] & G4[answer G] G1 & G2 & G3 & G4 --> VER[verifier<br/>checks each answer] VER --> NORM[normalise within the group<br/>mean and std] NORM --> AD2[advantage per answer] end

Reading it: both pipelines end in an advantage per answer, which feeds the same clipped update. PPO gets its baseline from a value network that must be trained, stored and run, which is roughly a second copy of the model. GRPO instead spends that compute on more answers per prompt and lets them grade each other. The verifier box can be any scorer, but GRPO shines when it is a program that checks the answer.

Verifiable rewards, and why GRPO trains reasoning

A verifiable reward comes from a check that can't be argued with: does the arithmetic equal 7, do the unit tests pass, does the proof check. No learned reward model, so no learned blind spots (see reward hacking, next). Here is GRPO on a toy "language model" that answers eight addition prompts with a single digit token. It starts out about 23% accurate and leans towards off-by-one mistakes.

GRPO on eight addition prompts: accuracy climbs from 23% to 99% in 60 steps, while the share of prompts whose whole group agreed, and so taught nothing, rises from a few percent to over 80%

Reading it: the blue line is the model's average chance of answering correctly, which climbs from 0.23 to 0.99 in 60 steps of 8 answers per prompt. The grey bars are the share of prompts whose group of 8 answers were all right or all wrong: a few percent at the start, over 80% by the end. They grow as the model masters the prompts: once every answer is right, the advantages are all zero and that prompt has nothing left to teach. Real GRPO training fights exactly this by filtering for prompts at the edge of the model's ability.

Why it matters in practice. This recipe, a verifier on the final answer plus GRPO, is how DeepSeek-R1-Zero learned to reason: rewarded only for correct final answers (and a required format), the model learned by itself to write longer chains of thought, to check its work and to back up from mistakes, because those behaviours raised the chance of a correct final answer. Each token in a long chain of thought shares its answer's advantage, so the whole chain is reinforced or discouraged together. See primer.ml.reasoning for what that training produces.

In code: group_advantages is the formula, verify is the checker, pretrained_logits is the weak starting model, and train_grpo runs the loop.

Reward hacking: the score is not the goal

Everyday picture. A school pays tutors by the number of pages of homework feedback they write. Feedback gets longer, not better. An RL agent is the most literal-minded employee imaginable: it optimises the number you wrote down, not the thing you meant. When those two differ, it finds the difference. This is reward hacking (also called specification gaming), and it is Goodhart's law in code: when a measure becomes a target, it ceases to be a good measure.

Tiny worked example. The goal is a good answer, and good answers here are about 4 sentences long. People rated answers of 1 to 4 sentences, and on those, longer really was better. A reward model fitted to their ratings learns a straight line: "each sentence is worth 0.19 points". Nobody ever rated a 10-sentence answer, so nobody told the reward model it was bad:

Sentences 1 2 3 4 6 8 10
True quality 0.44 0.75 0.94 1.00 0.75 0.00 −1.25
Reward model 0.50 0.69 0.88 1.06 1.44 1.81 2.19

The reward model's favourite answer, 10 sentences, is the worst one.

Level 3: the formula and its symbols

$$ \hat r(n) = w_0 + w_1 n, \qquad w_1 = \frac{\sum_i (n_i - \bar n)(q_i - \bar q)}{\sum_i (n_i - \bar n)^2}, \qquad w_0 = \bar q - w_1 \bar n $$

Symbols

Symbol Meaning here In the example
$n$ the answer's length in sentences 1 … 10
$\hat r(n)$ the reward model's score for a length-$n$ answer (the hat marks an estimate) 2.19 at n = 10
$n_i, q_i$ the rated examples: a length and its true quality (1, 0.44), …, (4, 1.0)
$\bar n, \bar q$ their averages (the bar means "mean") 2.5, 0.781
$w_1$ the fitted slope: points per extra sentence 0.1875
$w_0$ the fitted intercept 0.3125

In words: "the reward model is the straight line that best fits the ratings it saw: its slope is how length and quality moved together in the data, and it passes through the average point."

With the numbers: the deviations of length are (−1.5, −0.5, 0.5, 1.5) and of quality (−0.344, −0.031, 0.156, 0.219). Their products add to 0.9375 and the squared length deviations to 5, so w₁ = 0.1875 and w₀ = 0.781 − 0.1875 × 2.5 = 0.3125. At n = 10 it scores 2.19.

Level 3: in Python

In Python:

n = [1, 2, 3, 4]
q = [1 - ((n_i - 4) / 4) ** 2 for n_i in n]
q  # → [0.4375, 0.75, 0.9375, 1.0]
n_bar, q_bar = sum(n) / 4, sum(q) / 4
w1 = sum((a - n_bar) * (b - q_bar) for a, b in zip(n, q)) / sum((a - n_bar) ** 2 for a in n)
w0 = q_bar - w1 * n_bar
(w0, w1)  # → (0.3125, 0.1875)
# the reward model's score for a 10-sentence answer, and its true quality
(w0 + w1 * 10, 1 - ((10 - 4) / 4) ** 2)  # → (2.1875, -1.25)

True quality peaks at 4 sentences and falls below zero past 8, while the reward model's straight line, fitted on lengths 1 to 4, keeps rising to 2.19 at 10 sentences

Reading it: the shaded strip is the only region anyone rated. Inside it, the reward model (red) and the truth (blue) agree on the direction: longer is better. Outside it, the reward model is extrapolating a straight line into territory it has never seen, while the truth turns over and dives. Every learned reward model has regions like this, and an optimiser is a machine for finding them.

flowchart LR GOAL[What we want<br/>helpful answers] -->|people rate a sample| DATA[Ratings] DATA -->|fit| RM[Reward model<br/>a proxy for the goal] RM -->|reward| OPT[RL optimiser] OPT --> POL[Policy] POL -->|drifts to where<br/>proxy and goal disagree| GAP[Gap:<br/>high reward, low quality] KL[KL leash] -.->|limits drift| POL VER[Verifiable reward] -.->|no learned gap| OPT

Reading it: the solid path is how a reward is usually made: the goal is sampled by people, the ratings train a model, and that model, not the goal, is what the optimiser sees. The optimiser pushes the policy wherever the reward is highest, which, once the easy gains are taken, is wherever the proxy is most wrong. The dotted arrows are the two main defences: a leash that limits how far the policy can drift from where the ratings were collected, and a reward that has no learned gap to exploit.

Now train the starting model (which writes 2 or 3 sentences, a little too short) against each reward:

Over 300 steps, optimising the flawed reward first raises true quality to 0.93 then drives it down to 0.75, below where it started, while a KL leash holds it near 0.84 and the verifiable reward climbs to 1.0; the right panel shows the best leashed policy's true quality peaking at a moderate KL and collapsing as the leash loosens

Reading it: on the left, true quality over 300 training steps. Against the flawed reward with no leash (red), quality rises at first, from 0.80 to 0.93, because lengthening a too-short answer genuinely helps. It rests on 5-sentence answers for a while, then the reward model's pull wins again: around step 150 the policy jumps to 6 sentences and quality falls to 0.75, below where it started, while the reward model's score keeps climbing. That rise-then-fall is the signature of reward hacking, and it is why "the reward went up" proves nothing. With a KL leash (β = 0.3, orange) the policy also overshoots a little, but the leash holds it at 0.84. Against the true, verifiable quality (blue) it reaches 1.0. On the right, the best policy for each leash strength, placed by how far it strays from the reference (its KL). Near zero KL it barely moves; at moderate KL true quality peaks; as the leash loosens further, the proxy score keeps rising and the true quality collapses below zero.

The right-hand panel uses a closed form: the best policy under a KL leash has an exact formula.

Level 3: the formula and its symbols

$$ \pi^*(a) = \frac{\pi_{\text{ref}}(a)\, e^{R(a)/\beta}}{Z}, \qquad Z = \sum_{b} \pi_{\text{ref}}(b)\, e^{R(b)/\beta} $$

Symbols

Symbol Meaning here In the example
$\pi^*(a)$ the policy that maximises $\mathbb{E}[R] - \beta\,\mathrm{KL}(\pi \,|\, \pi_{\text{ref}})$ (0.731, 0.269)
$\pi_{\text{ref}}(a)$ the reference model's probability for action $a$ (0.5, 0.5)
$R(a)$ the reward for action $a$ (1, 0)
$\beta$ leash strength 1
$e^{R(a)/\beta}$ a boost that grows with reward; a small β makes it enormous $e^1 = 2.718$
$b$ a counter over every action, so $Z$ adds up all of them
$Z$ the total, so the probabilities add to 1 1.859

In words: "the best leashed policy starts from the reference and multiplies each action's probability by e to the power reward over β, then rescales so everything adds to 1."

With the numbers: two actions, reference (0.5, 0.5), rewards (1, 0), β = 1: weights 0.5 × 2.718 = 1.359 and 0.5 × 1 = 0.5, total 1.859, so π* = (0.731, 0.269). At β = 0.1 the boost is e¹⁰ ≈ 22,026 and π* puts 99.995% on the rewarded action; as β grows, π* returns to the reference.

Level 3: in Python

In Python:

import math
ref, R = [0.5, 0.5], [1.0, 0.0]
def best_policy(beta):
    w = [p * math.exp(r / beta) for p, r in zip(ref, R)]
    return [round(w_a / sum(w), 5) for w_a in w]
best_policy(1.0)  # → [0.73106, 0.26894]
best_policy(0.1)  # → [0.99995, 5e-05]
best_policy(100.0)  # → [0.5025, 0.4975]

This is the same formula DPO starts from (primer.ml.training_stages): it is why β means the same thing in RLHF and DPO.

Why it matters in practice. Reward hacking shows up wherever RL does. A boat-racing game agent that learned to circle forever collecting bonus targets instead of finishing the race. RLHF'd chat models that learned long, flattering answers score well with raters (length bias and sycophancy). Coding models rewarded for passing tests that learned to edit or special-case the tests. The defences, strongest first:

  1. Verifiable rewards where the task allows: run the tests, check the answer. A check has no learned blind spot (though a buggy check does).
  2. Better reward models: rate the policy's current outputs and retrain, so the reward model sees the regions the policy is exploring; use ensembles; penalise known exploits such as length directly.
  3. A KL leash to keep the policy near the data the reward was fit on.
  4. Watch a held-out measure of the real goal (human review, a separate evaluation set, see primer.agents.evals) and stop when it turns down, even if the reward is still climbing.

In code: true_quality is the goal, fit_reward_model and proxy_reward are the flawed reward, optimise_lengths trains against either with an optional leash, and kl_regularised_optimum with leash_sweep gives the closed-form best policy for each β.

In 20 seconds

  • RL learns from a score, not an answer: sample an action from the policy, get a reward, make rewarded actions more likely.
  • REINFORCE: step along reward × ∇ log π(action); for softmax, ∇ log π is one-hot minus the probabilities. Unbiased but noisy.
  • Baselines subtract the typical reward, turning rewards into advantages. Same average gradient, far less variance.
  • PPO reuses each batch for several steps, clips the probability ratio to 1 ± ε so no step goes too far, and uses a value network as baseline plus a KL leash to a reference model.
  • GRPO drops the value network: sample a group of answers per prompt and normalise rewards within the group. With a verifier as reward, it is how reasoning models are trained.
  • Reward hacking: the policy optimises the reward you wrote, not the goal you meant. Defend with verifiable rewards, better reward models, a KL leash and a held-out check of the real goal.

Self-test questions

How does reinforcement learning differ from supervised learning? Supervised learning is given the correct output for every input and learns to copy it. Reinforcement learning is given only a score for the output it produced, so it must try things, see how they score and shift probability towards what scored well. It fits tasks where judging an answer is easy but writing the perfect one is not.

What is the log-probability trick, and why is it needed? The gradient of expected reward, Σ ∇π(a) R(a), can't be computed without knowing every action's reward. Rewriting ∇π = π ∇log π turns it into an average over actions sampled from the policy, E[R ∇log π(a)], so each sampled action and its reward give an unbiased estimate of the gradient.

Why does subtracting a baseline not change the expected gradient? Because E[b ∇log π(a)] = b ∇ Σ π(a) = b ∇ 1 = 0: probabilities always add to 1, so the baseline's push averages to nothing. It only removes the noise that comes from rewards being large or all the same sign.

In PPO, what is the probability ratio, and what does clipping it do? The ratio is the current policy's probability of a sampled action divided by the probability under the policy that sampled it; it measures how far the policy has moved on that sample. Clipping stops counting movement beyond 1 ± ε in the direction the advantage favours, so reusing a batch for several steps can't push the policy far from where the data came from.

Why does PPO's objective take the minimum of the clipped and unclipped terms? To stay pessimistic. Gains are capped once the ratio leaves the band, but if a step made a bad action more likely, the full penalty still applies, so the objective never rewards a harmful move.

How does GRPO get a baseline without a value network? It samples several answers to the same prompt and uses their mean reward as the baseline, dividing by their standard deviation to set the scale. Each answer is judged against its siblings, which saves training and serving a second model the size of the policy.

What happens in GRPO when every answer in a group gets the same reward? Every advantage is zero, so that prompt contributes no gradient. Prompts that are always solved or never solved teach nothing; learning comes from prompts at the edge of the model's ability.

What is reward hacking, and why does a KL penalty help against it? Reward hacking is the policy maximising the reward as written while the real goal gets worse, usually by finding inputs where a learned reward model is wrong. The KL penalty charges the policy for drifting from the reference model, which keeps it near the kind of outputs the reward model was trained on, where the reward is still trustworthy.

Why are verifiable rewards attractive for training reasoning? A program that checks the final answer (or runs the tests) has no learned blind spots to exploit and costs nothing to label, so RL can run for a long time against it without the reward drifting away from correctness.

The papers behind this lesson

Further reading

on GitHub
   1r"""
   2# Reinforcement learning: learning from a score instead of an answer
   3
   4Run: `python -m primer.ml.reinforcement`
   5
   6## Level 1: The practitioner's guide
   7
   8**In one sentence.** Reinforcement learning trains a model from a score on
   9what it produced rather than from a correct answer to copy, which is how a
  10language model learns things nobody can write down (be helpful, reason to
  11the right answer) and how it learns to game a score that was written down
  12badly.
  13
  14**When you need it.** You need RL when you can judge an answer but cannot
  15write the perfect one: a proof either checks or it doesn't, tests pass or
  16fail, one reply is better than another though neither is "the" reply.
  17The tell: you find yourself writing a grader, a checker or a rubric instead
  18of example answers. You don't need it when you can write the answers
  19(supervised fine-tuning copies them, at a fraction of the cost) or when
  20you have pairs of better and worse answers and nothing more (DPO in
  21`primer.ml.training_stages` learns from pairs with no sampling loop). And
  22you never need to run the vendor's own RL: the helpfulness, harmlessness and
  23reasoning of a hosted model were trained this way before you arrived, which
  24is why this lesson matters even if you never train: it explains why models
  25answer at length, flatter, and sometimes optimise the letter of your
  26instruction instead of its spirit.
  27
  28**Your options.** From the cheapest to the most committed:
  29
  30| Option | What it does | What it guarantees | What it costs | Where it lives |
  31|---|---|---|---|---|
  32| Supervised fine-tuning on demonstrations | Copy correct answers you wrote | The behaviour in the examples, nothing beyond them | Writing the answers | Your training stack, or a hosted API |
  33| DPO on preference pairs | Learn from "this one beat that one", no sampling during training | Shifts tone and choices without a reward model or an RL loop | Thousands of comparisons | Your training stack, or a hosted API |
  34| Hosted reinforcement fine-tuning with your grader | The vendor samples answers and reinforces the ones your grader scores high | An RL loop you don't build; the grader is still yours to get right | Grader design, many sampled answers per prompt, the vendor's price | The vendor's API (as of October 2026, OpenAI's fine-tuning platform no longer accepts new users, so check availability first) |
  35| GRPO with a verifiable reward | Sample a group of answers per prompt, check each, reinforce the above-average ones | A reward with no learned blind spot, and no second model to train | Generation dominates: 8 answers per prompt is a common default; a checker that cannot be argued with | Your training stack |
  36| PPO with a learned reward model | Train a reward model on ratings, then a value network and the policy against it, on a KL leash | Optimises a goal no program can check, such as helpfulness | Two extra models (a reward model and a value network; InstructGPT used 6B for both at every policy size), and the reward model's blind spots to defend against | Your training stack |
  37
  38**How to choose.** Ask what can judge an answer, and how much you trust it.
  39
  40- A program can check the final answer (arithmetic, unit tests, a format,
  41  a proof checker): GRPO with that check as the reward. In this lesson's
  42  toy, accuracy on eight addition prompts goes from 23% to 99% in 60 steps
  43  of 8 answers each; this recipe is how DeepSeek-R1-Zero learned to reason
  44  from rule-based rewards: a correct final answer and a required format.
  45- Only people can judge, and you have their ratings: a learned reward model
  46  with PPO, a KL leash, and a held-out measure of the real goal that you
  47  watch more closely than the reward.
  48- You have pairs but no budget for sampling: DPO.
  49- You can write the answers: supervised fine-tuning, and stop there.
  50- Whatever you pick, the policy optimises the reward you wrote, not the goal
  51  you meant. Before training, ask what a literal-minded optimiser would do
  52  with your reward, and measure the goal separately.
  53
  54**What it costs.** Sampling is the bill: every training example is a full
  55generation, and GRPO multiplies it by the group size, which is why PPO
  56reuses each batch for several passes and why training loops run a fast
  57inference engine beside the trainer. A learned baseline costs a second
  58model (PPO's value network); GRPO replaces it with the group's own mean,
  59which is why it was introduced as a way to cut PPO's memory. Noise costs
  60steps: in the lesson's two-arm toy the gradient estimate's variance is
  6130.25 without a baseline and 0 with one, and with rewards offset by five
  62points, 60% of training runs without a baseline lock onto the wrong arm
  63against 0% with one. Reusing a batch too hard costs calibration: 50
  64passes over 16 pulls push one arm to 0.84 unclipped, 0.41 with PPO's clip
  65at 0.2. And a KL leash costs a little reward on every answer (1.0 becomes
  660.931 for an answer whose probability doubled, at β = 0.1) to buy fluency
  67and safety from drift.
  68
  69**What breaks.**
  70
  71- **Reward hacking.** The policy finds where the reward and the goal
  72  disagree. In the toy, a reward model fitted on answers of 1 to 4
  73  sentences rates a 10-sentence answer 2.19, the worst answer of all; true
  74  quality rises from 0.80 to 0.93 and then falls to 0.75, below where it
  75  started, while the reward keeps climbing. Gao, Schulman and Hilton (2022)
  76  measured the same rise and fall at scale. "The reward went up" proves
  77  nothing; keep a held-out measure of the goal.
  78- **Length bias and sycophancy.** Raters prefer long, flattering
  79  answers, so the reward model does too, so the model becomes that.
  80  Penalise length directly and rate the policy's current outputs, not
  81  stale ones.
  82- **Gaming the checker.** A coding model rewarded for passing tests learns
  83  to edit or special-case the tests. A verifier has no learned blind spot,
  84  but a buggy or bypassable one is a reward model with extra steps.
  85- **Groups with nothing to teach.** When every answer in a GRPO group is
  86  right or all are wrong, every advantage is zero: by the end of the toy
  87  run, 84% of groups are unanimous. Filter for prompts at the edge of the
  88  model's ability.
  89- **Too loose a leash.** As β falls, the best policy piles onto whatever the
  90  reward likes (99.995% on one action at β = 0.1 in the toy) and true
  91  quality collapses; too tight and nothing moves. Sweep it.
  92- **Collapsed exploration.** An arm whose probability hits zero is never
  93  tried again, so the policy can lock onto a mediocre answer early.
  94
  95**In the wild.** InstructGPT (Ouyang et al., 2022) set the pattern of PPO
  96against a learned reward model with a per-token KL penalty, and every chat
  97assistant since inherits its habits. DeepSeekMath (Shao et al., 2024)
  98introduced GRPO, reaching 51.7% on the MATH benchmark with a 7-billion-parameter model; DeepSeek-R1 (2025) trained
  99reasoning with RL on rule-based rewards and no human-written reasoning
 100traces. Hugging Face TRL's GRPOTrainer takes reward functions as plain
 101Python callables or a reward model, samples 8 generations per prompt by
 102default, and can generate with vLLM; OpenAI's model optimization guide
 103lists reinforcement fine-tuning, where you supply the grader (as of October 2026
 104OpenAI's fine-tuning platform no longer accepts new users). Sutton and
 105Barto's textbook and OpenAI's Spinning Up are the standard longer reads.
 106
 107**Go deeper.** Level 2 builds it all on a three-armed slot machine:
 108REINFORCE as one line of arithmetic, why a baseline removes noise without
 109bias, PPO's ratio and clip on a table of four cases, the KL leash as a fee
 110per answer, GRPO's advantages from a group of four, and a reward model that
 111loves length, trained against until quality falls, each with a figure you
 112can rerun. If you only needed to choose a training signal, you are done.
 113
 114## Level 2: How it works, from scratch
 115
 116Think of teaching a dog to sit. You can't show it the right answer: you can
 117only wait for it to try something and give it a treat when the something was
 118good. Over many tries the dog does more of what earned treats and less of
 119what didn't. Nobody ever told it what "sit" means; it worked it out from a
 120score.
 121
 122That is **reinforcement learning (RL)**. An **agent** (the dog, or a
 123language model) takes an **action** (sits, or writes an answer), the world
 124hands back a **reward** (a treat, or a score from a grader), and the agent
 125adjusts itself so that rewarded actions become more likely. The agent's
 126current habits, written as a probability for every action, are called its
 127**policy**.
 128
 129Compare this with ordinary supervised learning (`primer.ml.losses`), where
 130every example comes with the correct answer attached. In RL there is no
 131answer key, only a score after the fact. That is exactly the situation a
 132language model is in once pretraining is over: for "prove this theorem" or
 133"write a helpful reply" there is no single correct text to copy, but a
 134checker or a judge can say how good an attempt was. RL is how the model
 135learns from those judgements (`primer.ml.training_stages` places it in the
 136training pipeline).
 137
 138```mermaid
 139flowchart LR
 140  P[Policy<br/>a probability for every action] -->|sample| A[Action]
 141  A --> E[Environment<br/>a slot machine, a grader, a user]
 142  E --> R[Reward<br/>one number]
 143  R -->|nudge the policy| P
 144```
 145
 146**Reading it:** follow the loop clockwise. The policy is a set of
 147probabilities, and the action is *sampled* from it, so the agent sometimes
 148tries things it isn't sure about. The environment is whatever judges the
 149action; the agent can't see inside it. The only thing that comes back is one
 150number, the reward, and the only thing the agent can do with it is nudge its
 151own probabilities. Everything in this lesson is a better answer to one
 152question: how exactly should that nudge be computed?
 153
 154## A tiny worked example: three slot machines
 155
 156The simplest RL problem is a row of slot machines, called a **multi-armed
 157bandit** (a slot machine is a "one-armed bandit"). Machine A pays out 20% of
 158the time, B 50% and C 80%, but the agent isn't told that. Each pull pays 1
 159or 0. The agent must find C by pulling and seeing what happens.
 160
 161A policy here is three probabilities, one per arm. A policy that picks each
 162arm a third of the time earns, on average, (0.2 + 0.5 + 0.8) / 3 = **0.5**
 163per pull. A policy that always picks C earns **0.8**. Learning means moving
 164from the first policy to the second using nothing but the 1s and 0s.
 165
 166This one number, the average reward a policy expects, is what RL maximises:
 167
 168$$
 169J(\theta) = \sum_{a} \pi_\theta(a)\, R(a)
 170$$
 171
 172**Symbols**
 173
 174| Symbol | Meaning here | In the example |
 175|---|---|---|
 176| $a$ | one action: which arm to pull | A, B or C |
 177| $\theta$ | the policy's adjustable numbers (here, one score per arm, called **logits**) | (0, 0, 0) |
 178| $\pi_\theta(a)$ | the policy: the probability of picking action $a$, given $\theta$. Here softmax of the logits | 1/3 each |
 179| $R(a)$ | the average reward action $a$ pays (unknown to the agent) | 0.2, 0.5, 0.8 |
 180| $\sum_{a}$ | add up over every action | three terms |
 181| $J(\theta)$ | the expected reward: what the policy earns per pull, on average | 0.5 |
 182
 183**In words:** "the expected reward is each action's probability times its
 184average payout, added up over the actions."
 185
 186**With the numbers:** J = ⅓·0.2 + ⅓·0.5 + ⅓·0.8 = 0.5 for the uniform
 187policy, and 1·0.8 = 0.8 for the policy that always pulls C.
 188
 189**In Python:**
 190
 191```python
 192policy = [1/3, 1/3, 1/3]
 193# average payout of arms A, B, C
 194R = [0.2, 0.5, 0.8]
 195# J = Σ_a π(a) R(a)
 196round(sum(p * r for p, r in zip(policy, R)), 3)  # → 0.5
 197round(sum(p * r for p, r in zip([0, 0, 1], R)), 3)  # → 0.8
 198```
 199
 200The agent can't compute J, because it doesn't know R. It can only sample:
 201pull an arm, see a 1 or a 0. Every method below turns those samples into an
 202estimate of which way to move θ to make J bigger. Trying an arm you're
 203unsure of is called **exploration**; sticking with the best arm so far is
 204**exploitation**. A sampled policy does some of both automatically, as
 205long as no arm's probability has collapsed to zero.
 206
 207**In code:** `Bandit` hides the win chances and pays out one pull at a time; `expected_reward` is the formula above.
 208
 209## Policy gradients: do more of what worked (REINFORCE)
 210
 211**Everyday picture.** A football coach reviews the tape after a match. For
 212every play that led to a goal, they tell the team "a bit more of that"; for
 213plays that went nowhere, nothing. They don't need to know *why* the play
 214worked. Repeat over hundreds of matches and the team drifts towards the
 215plays that score.
 216
 217**Tiny worked example.** Start with logits (0, 0, 0), so each arm has
 218probability ⅓. The agent pulls C and wins: reward 1. The rule, explained
 219next, says: add to each logit *learning rate × reward × (1 if it's the
 220chosen arm, else 0, minus that arm's probability)*.
 221
 222| Arm | Chosen? | 1[chosen] − π | × reward 1 × rate 0.5 | New logit | New probability |
 223|---|---|---|---|---|---|
 224| A | no | 0 − ⅓ = −0.333 | −0.167 | −0.167 | **0.274** |
 225| B | no | 0 − ⅓ = −0.333 | −0.167 | −0.167 | **0.274** |
 226| C | yes | 1 − ⅓ = +0.667 | +0.333 | +0.333 | **0.452** |
 227
 228One lucky pull moved C from 33% to 45%. Had the pull paid 0, nothing would
 229have moved. Had the agent pulled A and won (A wins sometimes too), A would
 230have gone up instead. The rule is noisy, one pull at a time, but on average
 231the arm that wins most gets pushed up most.
 232
 233```mermaid
 234flowchart LR
 235  L[Logits θ] --> S[softmax<br/>probabilities π]
 236  S -->|sample| A[Action a]
 237  A --> ENV[Pull the arm] --> R[Reward R]
 238  S --> G["∇ log π(a)<br/>= one-hot(a) − π"]
 239  A --> G
 240  G --> M["× R × learning rate"]
 241  R --> M
 242  M -->|add to| L
 243```
 244
 245**Reading it:** the top path is acting: logits become probabilities, one
 246action is sampled, and the environment pays a reward. The lower path is
 247learning: from the action alone, work out which direction in logit space
 248makes that action more likely (the `∇ log π` box), then scale that direction
 249by how good the outcome was. A big reward is a big step towards repeating the
 250action; zero reward is no step.
 251
 252### The log-probability trick, decoded
 253
 254We want the **gradient** of J: for each logit, how much J rises if the
 255logit rises a little (see `primer.notation` for gradients from scratch).
 256The difficulty is that J is an average over actions we can only sample. The
 257trick rewrites the gradient as an average too, so a sample estimates it:
 258
 259$$
 260\nabla_\theta J(\theta) = \mathbb{E}_{a \sim \pi_\theta}\big[\, R(a)\, \nabla_\theta \log \pi_\theta(a) \,\big]
 261\qquad
 262\frac{\partial \log \pi_\theta(a)}{\partial z_k} = \mathbb{1}[k = a] - \pi_\theta(k)
 263$$
 264
 265**Symbols**
 266
 267| Symbol | Meaning here | In the example |
 268|---|---|---|
 269| $\nabla_\theta$ | "gradient with respect to θ": one slope per logit, collected into a vector | 3 slopes |
 270| $\mathbb{E}_{a \sim \pi_\theta}[\ldots]$ | **expected value**: the average of the bracket when $a$ is sampled from the policy | average over pulls |
 271| $\sim$ | "drawn from" | |
 272| $\log$ | the natural logarithm, the undo button for $e^x$ | $\log \tfrac13 = -1.10$ |
 273| $\nabla_\theta \log \pi_\theta(a)$ | the direction in logit space that makes action $a$ more likely, fastest | $(-\tfrac13, -\tfrac13, \tfrac23)$ for C |
 274| $z_k$ | the $k$-th logit (θ is the list of logits) | $z_3 = 0$ |
 275| $\partial$ | "partial derivative": the slope along one logit, holding the others still | |
 276| $\mathbb{1}[k = a]$ | 1 if $k$ is the chosen action, else 0 | (0, 0, 1) |
 277| $\pi_\theta(k)$ | the probability of action $k$ | ⅓ |
 278
 279**In words:** "the direction that raises expected reward is, on average,
 280the direction that makes the sampled action more likely, weighted by the
 281reward it earned. For a softmax policy, that direction is 'one for the
 282chosen action, minus every action's probability'."
 283
 284Why is this true? Because the slope of a probability equals the probability
 285times the slope of its log ($\nabla \pi = \pi \, \nabla \log \pi$, the chain
 286rule applied to log). So $\nabla J = \sum_a \nabla\pi(a) R(a) = \sum_a \pi(a)
 287\nabla\log\pi(a) R(a)$, and a sum weighted by $\pi(a)$ is an average over
 288samples from π. The update is then plain **gradient ascent**:
 289$\theta \leftarrow \theta + \alpha\, R\, \nabla_\theta \log \pi_\theta(a)$,
 290with learning rate $\alpha$ (`primer.ml.optimizers`).
 291
 292**With the numbers:** at the uniform policy the true gradient is
 293(−0.1, 0, +0.1): push C up, A down, leave B (which pays exactly the average)
 294alone. One sample, "pulled C, got 1", estimates it as 1 × (−⅓, −⅓, ⅔). The
 295step with α = 0.5 gives logits (−0.167, −0.167, 0.333) and probabilities
 296(0.274, 0.274, 0.452), as in the table.
 297
 298**In Python:**
 299
 300```python
 301import math
 302z = [0.0, 0.0, 0.0]
 303pi = [math.exp(z_k) / sum(math.exp(v) for v in z) for z_k in z]
 304# the true gradient: Σ_a π(a) R(a) (1[k=a] − π(k)), for each logit k
 305R = [0.2, 0.5, 0.8]
 306true_grad = [sum(pi[a] * R[a] * ((k == a) - pi[k]) for a in range(3)) for k in range(3)]
 307[round(g, 3) for g in true_grad]  # → [-0.1, 0.0, 0.1]
 308# one sample: pulled C (a = 2), reward 1
 309a, reward, alpha = 2, 1.0, 0.5
 310grad_log_pi = [(k == a) - pi[k] for k in range(3)]
 311[round(g, 3) for g in grad_log_pi]  # → [-0.333, -0.333, 0.667]
 312z = [z_k + alpha * reward * g for z_k, g in zip(z, grad_log_pi)]
 313[round(z_k, 3) for z_k in z]  # → [-0.167, -0.167, 0.333]
 314[round(math.exp(z_k) / sum(math.exp(v) for v in z), 3) for z_k in z]  # → [0.274, 0.274, 0.452]
 315```
 316
 317This algorithm is called **REINFORCE** (Williams, 1992). Run it for 500
 318pulls and the policy finds arm C:
 319
 320![REINFORCE on the three-armed bandit: the probability of arm C climbs from a third to about 0.96 within 500 pulls, and the average reward rises from 0.5 towards 0.8](figures/primer.ml.reinforcement.bandit_learning.svg)
 321
 322**Reading it:** on the left, each line is one arm's probability over 500
 323pulls. All three start at ⅓. C's line (the 80% arm) climbs towards 1 while A
 324and B sink; the wiggles are single lucky or unlucky pulls. On the right is
 325the reward, averaged over the last 50 pulls. It starts near 0.5 (random
 326pulling) and rises towards the dashed line at 0.8, the most any policy can
 327earn. Nobody told the agent which arm was best: the 1s and 0s were enough.
 328
 329**Why it matters in practice.** A language model is exactly this kind of
 330policy, with a vocabulary of tokens as its arms, and one sampled answer is a
 331string of sampled tokens. REINFORCE applies unchanged: sum the log
 332probabilities of every token in the answer, and scale the gradient by the
 333answer's reward. Every method below (PPO, GRPO) is REINFORCE with repairs.
 334
 335**In code:** `grad_log_prob` is one-hot minus the probabilities, `reinforce_step` is one update, `worked_reinforce_step` is the table above, and `train_reinforce` runs the whole loop against a `Bandit`.
 336
 337## Variance and baselines: grade on a curve
 338
 339**Everyday picture.** A teacher whose class all scores between 90 and 100
 340learns nothing by being told "you got a 92". What matters is whether 92 is
 341above or below the class average. Raw scores that are all large and
 342positive make every attempt look good; only the difference from typical
 343tells you which way to go.
 344
 345**Tiny worked example.** Two arms, a 50/50 policy, and every pull pays a lot:
 346arm 1 always pays 10, arm 2 always pays 12. Arm 2 is better, so the logit of
 347arm 2 should rise. Look at the REINFORCE estimate for that logit:
 348
 349| Pulled | Reward | 1[arm 2] − π(arm 2) | Estimate (no baseline) | Estimate (baseline 11) |
 350|---|---|---|---|---|
 351| arm 1 | 10 | 0 − 0.5 = −0.5 | 10 × −0.5 = **−5** | (10 − 11) × −0.5 = **+0.5** |
 352| arm 2 | 12 | 1 − 0.5 = +0.5 | 12 × +0.5 = **+6** | (12 − 11) × +0.5 = **+0.5** |
 353
 354Without a baseline the estimate is −5 or +6 depending on the coin flip.
 355It averages to +0.5, the right answer, but any single sample points the
 356wrong way half the time, and violently. Subtract the average reward, 11,
 357first and *every* sample says +0.5. Same average, no noise at all.
 358
 359$$
 360\nabla_\theta J(\theta) = \mathbb{E}_{a \sim \pi_\theta}\big[\, (R(a) - b)\, \nabla_\theta \log \pi_\theta(a) \,\big],
 361\qquad A(a) = R(a) - b
 362$$
 363
 364**Symbols**
 365
 366| Symbol | Meaning here | In the example |
 367|---|---|---|
 368| $b$ | the **baseline**: any number that doesn't depend on which action was taken; usually the average reward | 11 |
 369| $A(a)$ | the **advantage**: how much better action $a$ did than typical | −1 for arm 1, +1 for arm 2 |
 370| everything else | as in the REINFORCE formula above | |
 371
 372**In words:** "scale each step by how much better than typical the action
 373did, not by its raw reward."
 374
 375Why is it allowed? Because the baseline's contribution averages to zero:
 376$\mathbb{E}[\, b\, \nabla \log \pi(a)] = b \sum_a \nabla \pi(a) = b\, \nabla
 377\sum_a \pi(a) = b\, \nabla 1 = 0$. Probabilities always add to 1, so pushing
 378all of them up is impossible; the baseline only removes noise, never
 379signal.
 380
 381**With the numbers:** without a baseline the estimate's **variance** (the
 382average squared distance from its mean, see `primer.notation`) is
 383(25 + 36)/2 − 0.5² = **30.25**. With b = 11 it is **0**.
 384
 385**In Python:**
 386
 387```python
 388# arm 1 pays 10, arm 2 pays 12; the 1[arm 2] − π(arm 2) factor for each pull
 389pulls = [(10, -0.5), (12, +0.5)]
 390def mean_and_variance(b):
 391    estimates = [(reward - b) * direction for reward, direction in pulls]
 392    mean = sum(estimates) / 2
 393    return mean, sum((e - mean) ** 2 for e in estimates) / 2
 394mean_and_variance(b=0)  # → (0.5, 30.25)
 395mean_and_variance(b=11)  # → (0.5, 0.0)
 396```
 397
 398```mermaid
 399flowchart LR
 400  R[Reward R] --> MINUS["R − b"]
 401  B["Baseline b<br/>average reward so far"] --> MINUS
 402  MINUS --> ADV{Advantage A}
 403  ADV -->|positive: better than typical| UP[make the action<br/>more likely]
 404  ADV -->|negative: worse than typical| DOWN[make the action<br/>less likely]
 405```
 406
 407**Reading it:** the baseline sits between the reward and the update. Its
 408job is to turn "how good was this?" into "how much better than usual was
 409this?". The sign of the advantage now decides the direction of the step,
 410so a below-average action is actively pushed down, even though its raw
 411reward was positive.
 412
 413To see it matter, give every arm of our bandit 5 extra points: rewards are
 414now 5 or 6 instead of 0 or 1, and nothing about which arm is best has
 415changed.
 416
 417![With rewards offset by 5, the gradient estimate's variance falls from about 20 to 0.15 with a baseline, and 20 training runs all find arm C with it, while without it most runs lock onto the wrong arm](figures/primer.ml.reinforcement.baselines.svg)
 418
 419**Reading it:** on the left, the variance of a one-pull gradient estimate
 420at the starting policy, on a log scale: about 20 without a baseline and
 421about 0.15 with one, over a hundred times smaller. On the right, each dot is
 422one of 20 training runs (400 pulls each), placed at its final probability
 423of picking C. With a running-average baseline (blue) every run ends near
 4240.95. Without one (red), the dots scatter to both ends: in most runs, early
 425pulls of a mediocre arm paid 5 and were pushed up hard, and the policy
 426committed before it ever learned C was better.
 427
 428**Why it matters in practice.** Every practical policy-gradient method uses
 429a baseline. PPO learns one with a second network, the **value network** (or
 430**critic**), which predicts the expected reward from each state. GRPO, below,
 431gets one for free by comparing several answers to the same prompt.
 432
 433**In code:** `gradient_estimate_stats` computes the exact mean and variance of the estimate, `sampled_gradient_variance` measures it from real pulls, and `train_reinforce` subtracts a running average when `baseline=True`.
 434
 435## PPO: take several steps, but never too far
 436
 437**Everyday picture.** A chef tests a new recipe on one evening's diners.
 438It would be wasteful to use their comments for just one small tweak, so the
 439chef makes several rounds of changes from the same comment cards. But the
 440further the recipe drifts from what the diners actually ate, the less their
 441comments apply, so the chef caps each change: never more than 20% more or
 442less of any ingredient per round.
 443
 444For a language model, sampling answers is the expensive part (every answer
 445is a full generation), so **PPO (Proximal Policy Optimization)** reuses each
 446batch of answers for several gradient steps. It needs a way to tell how far
 447the policy has moved since the batch was sampled, and a brake.
 448
 449**Tiny worked example.** The **probability ratio** compares the policy now
 450with the policy that generated the sample. A ratio of 1.5 means the current
 451policy is 50% more likely to produce that answer than when it was sampled.
 452With a clip range ε = 0.2 the ratio is allowed to count only between 0.8
 453and 1.2:
 454
 455| Ratio ρ | Advantage A | ρ·A | clip(ρ, 0.8, 1.2)·A | min of the two | What happened |
 456|---|---|---|---|---|---|
 457| 1.1 | +3 | 3.3 | 3.3 | **3.3** | inside the band: plain REINFORCE |
 458| 1.5 | +2 | 3.0 | 1.2 × 2 = 2.4 | **2.4** | good action already boosted enough: gain capped |
 459| 0.5 | −1 | −0.5 | 0.8 × −1 = −0.8 | **−0.8** | bad action already cut enough: capped |
 460| 1.5 | −1 | −1.5 | 1.2 × −1 = −1.2 | **−1.5** | bad action made *more* likely: full penalty |
 461
 462$$
 463\rho_t(\theta) = \frac{\pi_\theta(a_t \mid s_t)}{\pi_{\text{old}}(a_t \mid s_t)}
 464\qquad
 465L^{\text{CLIP}}(\theta) = \mathbb{E}_t\Big[\min\big(\rho_t A_t,\ \operatorname{clip}(\rho_t,\ 1-\varepsilon,\ 1+\varepsilon)\, A_t\big)\Big]
 466$$
 467
 468**Symbols**
 469
 470| Symbol | Meaning here | In the example |
 471|---|---|---|
 472| $t$ | one sample in the batch; for a language model, one token of one answer | row of the table |
 473| $s_t$ | the **state**: what the policy saw before acting (the prompt plus the tokens so far) | a prompt |
 474| $a_t$ | the action taken (the token generated) | an answer |
 475| $\pi_{\text{old}}$ | the policy as it was when the batch was sampled, frozen | |
 476| $\pi_\theta$ | the policy now, after some steps on this batch | |
 477| $\rho_t$ | the probability ratio, new over old | 1.5 |
 478| $A_t$ | the advantage of that sample | +2 |
 479| $\varepsilon$ | the clip range, typically 0.1 to 0.3 | 0.2 |
 480| $\operatorname{clip}(x, lo, hi)$ | $x$, but pushed back to $lo$ or $hi$ if it falls outside | clip(1.5, 0.8, 1.2) = 1.2 |
 481| $\min$ | the smaller of the two | min(3.0, 2.4) = 2.4 |
 482| $\mathbb{E}_t$ | the average over all samples in the batch | |
 483| $L^{\text{CLIP}}$ | the objective PPO climbs | |
 484
 485**In words:** "for each sample, take the ratio-weighted advantage, but once
 486the ratio has moved more than ε from 1 in the direction the advantage wants,
 487stop counting further movement; and always take the more pessimistic of the
 488clipped and unclipped versions."
 489
 490The slope of ρ·A is exactly REINFORCE's gradient scaled by ρ (because
 491$\nabla\rho = \rho\,\nabla\log\pi_\theta$), so inside the band PPO is
 492REINFORCE with importance weighting. Outside it, the clipped term is flat:
 493that sample stops pushing.
 494
 495**With the numbers:** the four rows of the table, computed:
 496
 497**In Python:**
 498
 499```python
 500def clip(x, lo, hi):
 501    return max(lo, min(x, hi))
 502def L_clip(rho, A, eps=0.2):
 503    return min(rho * A, clip(rho, 1 - eps, 1 + eps) * A)
 504[round(L_clip(rho, A), 2) for rho, A in [(1.1, 3), (1.5, 2), (0.5, -1), (1.5, -1)]]  # → [3.3, 2.4, -0.8, -1.5]
 505# past 1 + ε with a positive advantage, a higher ratio earns nothing more
 506L_clip(1.3, 2) == L_clip(1.6, 2)  # → True
 507```
 508
 509```mermaid
 510flowchart TB
 511  OLD[Policy π_old] -->|generate a batch| B[Answers + rewards]
 512  B --> V[Value network<br/>predicts expected reward]
 513  V --> ADV[Advantages A_t]
 514  B --> ADV
 515  ADV --> LOOP
 516  subgraph LOOP["Several passes over the same batch"]
 517    RATIO["ratio ρ = π_θ / π_old"] --> CLIP["clip to 1 ± ε, take the min"]
 518    CLIP --> KL["subtract β × KL to the reference model"]
 519    KL --> STEP[gradient step on θ]
 520    STEP --> RATIO
 521  end
 522  LOOP -->|π_θ becomes the new π_old| OLD
 523```
 524
 525**Reading it:** the outer loop is sampling, the expensive part: the frozen
 526old policy writes a batch of answers and a value network turns their rewards
 527into advantages. The inner loop reuses that batch for several passes. Each
 528pass recomputes how far the policy has moved (the ratio), stops counting
 529movement beyond the band (the clip), and applies the KL leash described
 530below. When the passes are done, the updated policy becomes the new sampler
 531and the cycle repeats.
 532
 533To see the clip work, take one batch of 16 pulls from the bandit (C won all
 534four of its pulls, A lost all four) and make 50 passes over it:
 535
 536![Two panels: the clipped objective is flat outside the 0.8 to 1.2 band on the side the advantage favours; over 50 passes the unclipped ratio climbs past 2.5 while the clipped one levels off near 1.24](figures/primer.ml.reinforcement.ppo_clip.svg)
 537
 538**Reading it:** on the left is the objective for one sample as its ratio
 539changes. For a positive advantage (blue), the line rises with the ratio
 540until 1.2, then goes flat: no reward for pushing further. For a negative
 541advantage (red), it goes flat below 0.8. The dashed lines are what
 542REINFORCE would keep climbing. On the right, the largest ratio in the batch
 543after each of 50 passes. Unclipped (red), the policy chases the same 16
 544pulls further every pass until C's probability is 2.5 times what it was,
 545about 0.84 from sixteen pulls, which is wildly overconfident. Clipped
 546(blue), it levels off near 1.24: the brake is not a hard wall (other
 547samples can still nudge the policy), but it removes the incentive to
 548overfit one batch.
 549
 550**In code:** `ppo_clipped_objective` is the formula, `clipped_policy_update` makes several passes over one batch with or without the clip, and `ppo_drift_experiment` is the 50-pass comparison.
 551
 552### The KL leash: stay close to where you started
 553
 554**Everyday picture.** A dog on a long leash can explore, but it can't run
 555off a cliff. In RL for language models, the leash ties the policy to a
 556frozen copy of the model it started from, the **reference model** (usually
 557the model after supervised fine-tuning).
 558
 559**Tiny worked example.** An answer earns reward 1.0. The policy now gives
 560it probability 0.6; the reference gave it 0.3. With leash strength β = 0.1,
 561the reward actually used for training is 1.0 − 0.1 × ln(0.6 / 0.3) =
 5621.0 − 0.1 × 0.693 = **0.931**. The policy pays a small fee for having
 563doubled that answer's probability.
 564
 565$$
 566R'(a) = R(a) - \beta \log\frac{\pi_\theta(a)}{\pi_{\text{ref}}(a)}
 567\qquad
 568\mathbb{E}_{a\sim\pi_\theta}\big[R'(a)\big] = \mathbb{E}_{a\sim\pi_\theta}\big[R(a)\big] - \beta\, \mathrm{KL}(\pi_\theta \,\|\, \pi_{\text{ref}})
 569$$
 570
 571**Symbols**
 572
 573| Symbol | Meaning here | In the example |
 574|---|---|---|
 575| $R(a)$ | the reward from the grader or reward model | 1.0 |
 576| $R'(a)$ | the reward after the leash's fee | 0.931 |
 577| $\beta$ | leash strength: how much a unit of drift costs | 0.1 |
 578| $\pi_{\text{ref}}(a)$ | the frozen reference model's probability for the answer | 0.3 |
 579| $\log\frac{\pi_\theta(a)}{\pi_{\text{ref}}(a)}$ | how much more (positive) or less (negative) likely the policy makes this answer than the reference | ln 2 = 0.693 |
 580| $\mathrm{KL}(\pi_\theta \,\|\, \pi_{\text{ref}})$ | **KL divergence**: the average of that log ratio over the policy's own answers; zero only when the two agree (decoded in `primer.ml.training_stages`) | |
 581
 582**In words:** "each answer's reward is docked in proportion to how much
 583more likely the policy has made it than the reference did; on average, that
 584fee is β times the KL divergence between the two."
 585
 586**With the numbers:** 1.0 − 0.1 × ln 2 = 0.931. Had the policy *halved* the
 587answer's probability instead (0.15), the log ratio would be −0.693 and the
 588reward would rise to 1.069: the leash pulls both ways.
 589
 590**In Python:**
 591
 592```python
 593import math
 594R, beta = 1.0, 0.1
 595pi, pi_ref = 0.6, 0.3
 596# R' = R − β log(π / π_ref)
 597round(R - beta * math.log(pi / pi_ref), 3)  # → 0.931
 598round(R - beta * math.log(0.15 / pi_ref), 3)  # → 1.069
 599```
 600
 601**Why it matters in practice.** This is the "penalty for drifting" in the
 602RLHF loop of `primer.ml.training_stages`, and the same β appears in DPO. It
 603stops the policy from forgetting fluent language while it chases reward, and
 604it is the first line of defence against reward hacking, below.
 605
 606**In code:** `kl_penalised_reward` is R′; `clipped_policy_update` adds the leash's gradient when given a reference policy and a β, and `primer.ml.training_stages.kl_divergence` computes the KL itself.
 607
 608## GRPO: compare answers to the same question
 609
 610**Everyday picture.** Instead of hiring an examiner to predict how hard
 611each exam question is, a teacher gives the same question to eight students
 612and marks each answer relative to the others on *that* question. On an easy
 613question, getting it right is expected and earns little credit; on a hard
 614one, the only right answer stands out.
 615
 616PPO's baseline comes from the value network, a second model, often as large as the
 617policy, that has to be trained alongside it. **GRPO (Group Relative Policy
 618Optimization)** throws the value network away. For each prompt it samples
 619a group of answers and uses the group's own average as the baseline.
 620
 621**Tiny worked example.** The prompt is "3 + 4 =". The model samples four
 622answers and a checker scores them 1 if the answer is 7, else 0.
 623
 624| Group rewards | Mean | Std | Advantages |
 625|---|---|---|---|
 626| 1, 0, 0, 1 | 0.5 | 0.5 | **+1, −1, −1, +1** |
 627| 1, 0, 0, 0 | 0.25 | 0.433 | **+1.73, −0.58, −0.58, −0.58** |
 628| 1, 1, 1, 1 | 1 | 0 | **0, 0, 0, 0** |
 629
 630A lone right answer in a mostly wrong group earns a big advantage: it's
 631rare, so it's strong evidence. A group that is all right (or all wrong)
 632earns nothing: there is no contrast, so there is nothing to learn from.
 633
 634$$
 635A_i = \frac{r_i - \operatorname{mean}(r_1, \ldots, r_G)}{\operatorname{std}(r_1, \ldots, r_G)}
 636$$
 637
 638**Symbols**
 639
 640| Symbol | Meaning here | In the example |
 641|---|---|---|
 642| $G$ | the group size: answers sampled per prompt | 4 |
 643| $i$ | which answer in the group | 1 … 4 |
 644| $r_i$ | the reward for answer $i$ | 1, 0, 0, 0 |
 645| $\operatorname{mean}(\ldots)$ | the group's average reward: the baseline | 0.25 |
 646| $\operatorname{std}(\ldots)$ | the group's **standard deviation**, the typical distance from the mean (square root of the variance) | 0.433 |
 647| $A_i$ | answer $i$'s advantage, shared by every token of that answer | +1.73 |
 648
 649**In words:** "an answer's advantage is how far its reward sits above the
 650group's average, measured in units of the group's spread."
 651
 652**With the numbers:** for (1, 0, 0, 0): mean 0.25, variance
 653(0.75² + 3 × 0.25²) / 4 = 0.1875, std √0.1875 = 0.433, so the right answer
 654gets 0.75 / 0.433 = 1.73 and each wrong one −0.25 / 0.433 = −0.58.
 655
 656**In Python:**
 657
 658```python
 659import statistics
 660def advantages(r):
 661    mu, sd = statistics.mean(r), statistics.pstdev(r)
 662    return [round((r_i - mu) / sd, 2) if sd else 0.0 for r_i in r]
 663advantages([1, 0, 0, 1])  # → [1.0, -1.0, -1.0, 1.0]
 664advantages([1, 0, 0, 0])  # → [1.73, -0.58, -0.58, -0.58]
 665advantages([1, 1, 1, 1])  # → [0.0, 0.0, 0.0, 0.0]
 666```
 667
 668The rest of GRPO is PPO: the same ratio, the same clip, the same KL leash to
 669a reference model, averaged over the group. (Some implementations divide by
 670the sample standard deviation, with G − 1, rather than the population one;
 671the idea is identical.)
 672
 673```mermaid
 674flowchart LR
 675  subgraph PPO["PPO"]
 676    P1[Prompt] --> A1[one answer]
 677    A1 --> RM1[reward]
 678    A1 --> VN[value network<br/>a second big model]
 679    RM1 --> AD1[advantage = reward − value]
 680    VN --> AD1
 681  end
 682  subgraph GRPO["GRPO"]
 683    P2[Prompt] --> G1[answer 1] & G2[answer 2] & G3[answer ...] & G4[answer G]
 684    G1 & G2 & G3 & G4 --> VER[verifier<br/>checks each answer]
 685    VER --> NORM[normalise within the group<br/>mean and std]
 686    NORM --> AD2[advantage per answer]
 687  end
 688```
 689
 690**Reading it:** both pipelines end in an advantage per answer, which feeds
 691the same clipped update. PPO gets its baseline from a value network that
 692must be trained, stored and run, which is roughly a second copy of the model.
 693GRPO instead spends that compute on more answers per prompt and lets them
 694grade each other. The verifier box can be any scorer, but GRPO shines when
 695it is a program that checks the answer.
 696
 697### Verifiable rewards, and why GRPO trains reasoning
 698
 699A **verifiable reward** comes from a check that can't be argued with: does
 700the arithmetic equal 7, do the unit tests pass, does the proof check. No
 701learned reward model, so no learned blind spots (see reward hacking, next).
 702Here is GRPO on a toy "language model" that answers eight addition prompts
 703with a single digit token. It starts out about 23% accurate and leans
 704towards off-by-one mistakes.
 705
 706![GRPO on eight addition prompts: accuracy climbs from 23% to 99% in 60 steps, while the share of prompts whose whole group agreed, and so taught nothing, rises from a few percent to over 80%](figures/primer.ml.reinforcement.grpo.svg)
 707
 708**Reading it:** the blue line is the model's average chance of answering
 709correctly, which climbs from 0.23 to 0.99 in 60 steps of 8 answers per
 710prompt. The grey bars are the share of prompts whose group of 8 answers
 711were all right or all wrong: a few percent at the start, over 80% by the
 712end. They grow as the model masters the prompts:
 713once every answer is right, the advantages are all zero and that prompt has
 714nothing left to teach. Real GRPO training fights exactly this by filtering
 715for prompts at the edge of the model's ability.
 716
 717**Why it matters in practice.** This recipe, a verifier on the final
 718answer plus GRPO, is how DeepSeek-R1-Zero learned to reason: rewarded only
 719for correct final answers (and a required format), the model learned by
 720itself to write longer chains of thought, to check its work and to back
 721up from mistakes, because those behaviours raised the chance of a correct
 722final answer. Each token in a long chain of thought shares its answer's
 723advantage, so the whole chain is reinforced or discouraged together. See
 724`primer.ml.reasoning` for what that training produces.
 725
 726**In code:** `group_advantages` is the formula, `verify` is the checker, `pretrained_logits` is the weak starting model, and `train_grpo` runs the loop.
 727
 728## Reward hacking: the score is not the goal
 729
 730**Everyday picture.** A school pays tutors by the number of pages of
 731homework feedback they write. Feedback gets longer, not better. An RL agent
 732is the most literal-minded employee imaginable: it optimises the number you
 733wrote down, not the thing you meant. When those two differ, it finds the
 734difference. This is **reward hacking** (also called **specification
 735gaming**), and it is Goodhart's law in code: *when a measure becomes a
 736target, it ceases to be a good measure.*
 737
 738**Tiny worked example.** The goal is a good answer, and good answers here
 739are about 4 sentences long. People rated answers of 1 to 4 sentences, and
 740on those, longer really was better. A reward model fitted to their ratings
 741learns a straight line: "each sentence is worth 0.19 points". Nobody ever
 742rated a 10-sentence answer, so nobody told the reward model it was bad:
 743
 744| Sentences | 1 | 2 | 3 | **4** | 6 | 8 | **10** |
 745|---|---|---|---|---|---|---|---|
 746| True quality | 0.44 | 0.75 | 0.94 | **1.00** | 0.75 | 0.00 | −1.25 |
 747| Reward model | 0.50 | 0.69 | 0.88 | 1.06 | 1.44 | 1.81 | **2.19** |
 748
 749The reward model's favourite answer, 10 sentences, is the worst one.
 750
 751$$
 752\hat r(n) = w_0 + w_1 n,
 753\qquad
 754w_1 = \frac{\sum_i (n_i - \bar n)(q_i - \bar q)}{\sum_i (n_i - \bar n)^2},
 755\qquad
 756w_0 = \bar q - w_1 \bar n
 757$$
 758
 759**Symbols**
 760
 761| Symbol | Meaning here | In the example |
 762|---|---|---|
 763| $n$ | the answer's length in sentences | 1 … 10 |
 764| $\hat r(n)$ | the reward model's score for a length-$n$ answer (the hat marks an estimate) | 2.19 at n = 10 |
 765| $n_i, q_i$ | the rated examples: a length and its true quality | (1, 0.44), …, (4, 1.0) |
 766| $\bar n, \bar q$ | their averages (the bar means "mean") | 2.5, 0.781 |
 767| $w_1$ | the fitted slope: points per extra sentence | 0.1875 |
 768| $w_0$ | the fitted intercept | 0.3125 |
 769
 770**In words:** "the reward model is the straight line that best fits the
 771ratings it saw: its slope is how length and quality moved together in the
 772data, and it passes through the average point."
 773
 774**With the numbers:** the deviations of length are (−1.5, −0.5, 0.5, 1.5)
 775and of quality (−0.344, −0.031, 0.156, 0.219). Their products add to 0.9375
 776and the squared length deviations to 5, so w₁ = 0.1875 and
 777w₀ = 0.781 − 0.1875 × 2.5 = 0.3125. At n = 10 it scores 2.19.
 778
 779**In Python:**
 780
 781```python
 782n = [1, 2, 3, 4]
 783q = [1 - ((n_i - 4) / 4) ** 2 for n_i in n]
 784q  # → [0.4375, 0.75, 0.9375, 1.0]
 785n_bar, q_bar = sum(n) / 4, sum(q) / 4
 786w1 = sum((a - n_bar) * (b - q_bar) for a, b in zip(n, q)) / sum((a - n_bar) ** 2 for a in n)
 787w0 = q_bar - w1 * n_bar
 788(w0, w1)  # → (0.3125, 0.1875)
 789# the reward model's score for a 10-sentence answer, and its true quality
 790(w0 + w1 * 10, 1 - ((10 - 4) / 4) ** 2)  # → (2.1875, -1.25)
 791```
 792
 793![True quality peaks at 4 sentences and falls below zero past 8, while the reward model's straight line, fitted on lengths 1 to 4, keeps rising to 2.19 at 10 sentences](figures/primer.ml.reinforcement.length_rewards.svg)
 794
 795**Reading it:** the shaded strip is the only region anyone rated. Inside
 796it, the reward model (red) and the truth (blue) agree on the direction:
 797longer is better. Outside it, the reward model is extrapolating a straight
 798line into territory it has never seen, while the truth turns over and
 799dives. Every learned reward model has regions like this, and an optimiser
 800is a machine for finding them.
 801
 802```mermaid
 803flowchart LR
 804  GOAL[What we want<br/>helpful answers] -->|people rate a sample| DATA[Ratings]
 805  DATA -->|fit| RM[Reward model<br/>a proxy for the goal]
 806  RM -->|reward| OPT[RL optimiser]
 807  OPT --> POL[Policy]
 808  POL -->|drifts to where<br/>proxy and goal disagree| GAP[Gap:<br/>high reward, low quality]
 809  KL[KL leash] -.->|limits drift| POL
 810  VER[Verifiable reward] -.->|no learned gap| OPT
 811```
 812
 813**Reading it:** the solid path is how a reward is usually made: the goal is
 814sampled by people, the ratings train a model, and that model, not the goal,
 815is what the optimiser sees. The optimiser pushes the policy wherever the
 816reward is highest, which, once the easy gains are taken, is wherever the
 817proxy is most wrong. The dotted arrows are the two main defences: a leash
 818that limits how far the policy can drift from where the ratings were
 819collected, and a reward that has no learned gap to exploit.
 820
 821Now train the starting model (which writes 2 or 3 sentences, a little too
 822short) against each reward:
 823
 824![Over 300 steps, optimising the flawed reward first raises true quality to 0.93 then drives it down to 0.75, below where it started, while a KL leash holds it near 0.84 and the verifiable reward climbs to 1.0; the right panel shows the best leashed policy's true quality peaking at a moderate KL and collapsing as the leash loosens](figures/primer.ml.reinforcement.reward_hacking.svg)
 825
 826**Reading it:** on the left, true quality over 300 training steps. Against
 827the flawed reward with no leash (red), quality *rises* at first, from 0.80
 828to 0.93, because lengthening a too-short answer genuinely helps. It rests
 829on 5-sentence answers for a while, then the reward model's pull wins again:
 830around step 150 the policy jumps to 6 sentences and quality falls to 0.75,
 831below where it started, while the reward model's score keeps climbing.
 832That rise-then-fall is the signature of reward hacking, and it is why "the
 833reward went up" proves nothing. With a KL leash (β = 0.3, orange) the
 834policy also overshoots a little, but the leash holds it at 0.84. Against the true, verifiable quality (blue) it
 835reaches 1.0. On the right, the best policy for each leash strength, placed
 836by how far it strays from the reference (its KL). Near zero KL it barely
 837moves; at moderate KL true quality peaks; as the leash loosens further, the
 838proxy score keeps rising and the true quality collapses below zero.
 839
 840The right-hand panel uses a closed form: the best policy under a KL leash
 841has an exact formula.
 842
 843$$
 844\pi^*(a) = \frac{\pi_{\text{ref}}(a)\, e^{R(a)/\beta}}{Z},
 845\qquad
 846Z = \sum_{b} \pi_{\text{ref}}(b)\, e^{R(b)/\beta}
 847$$
 848
 849**Symbols**
 850
 851| Symbol | Meaning here | In the example |
 852|---|---|---|
 853| $\pi^*(a)$ | the policy that maximises $\mathbb{E}[R] - \beta\,\mathrm{KL}(\pi \,\|\, \pi_{\text{ref}})$ | (0.731, 0.269) |
 854| $\pi_{\text{ref}}(a)$ | the reference model's probability for action $a$ | (0.5, 0.5) |
 855| $R(a)$ | the reward for action $a$ | (1, 0) |
 856| $\beta$ | leash strength | 1 |
 857| $e^{R(a)/\beta}$ | a boost that grows with reward; a small β makes it enormous | $e^1 = 2.718$ |
 858| $b$ | a counter over every action, so $Z$ adds up all of them | |
 859| $Z$ | the total, so the probabilities add to 1 | 1.859 |
 860
 861**In words:** "the best leashed policy starts from the reference and
 862multiplies each action's probability by e to the power reward over β, then
 863rescales so everything adds to 1."
 864
 865**With the numbers:** two actions, reference (0.5, 0.5), rewards (1, 0),
 866β = 1: weights 0.5 × 2.718 = 1.359 and 0.5 × 1 = 0.5, total 1.859, so
 867π* = (0.731, 0.269). At β = 0.1 the boost is e¹⁰ ≈ 22,026 and π* puts
 86899.995% on the rewarded action; as β grows, π* returns to the reference.
 869
 870**In Python:**
 871
 872```python
 873import math
 874ref, R = [0.5, 0.5], [1.0, 0.0]
 875def best_policy(beta):
 876    w = [p * math.exp(r / beta) for p, r in zip(ref, R)]
 877    return [round(w_a / sum(w), 5) for w_a in w]
 878best_policy(1.0)  # → [0.73106, 0.26894]
 879best_policy(0.1)  # → [0.99995, 5e-05]
 880best_policy(100.0)  # → [0.5025, 0.4975]
 881```
 882
 883This is the same formula DPO starts from (`primer.ml.training_stages`): it
 884is why β means the same thing in RLHF and DPO.
 885
 886**Why it matters in practice.** Reward hacking shows up wherever RL does.
 887A boat-racing game agent that learned to circle forever collecting bonus
 888targets instead of finishing the race. RLHF'd chat models that learned
 889long, flattering answers score well with raters (length bias and
 890sycophancy). Coding models rewarded for passing tests that learned to edit
 891or special-case the tests. The defences, strongest first:
 892
 8931. **Verifiable rewards** where the task allows: run the tests, check the
 894   answer. A check has no learned blind spot (though a buggy check does).
 8952. **Better reward models:** rate the policy's *current* outputs and
 896   retrain, so the reward model sees the regions the policy is exploring;
 897   use ensembles; penalise known exploits such as length directly.
 8983. **A KL leash** to keep the policy near the data the reward was fit on.
 8994. **Watch a held-out measure of the real goal** (human review, a separate
 900   evaluation set, see `primer.agents.evals`) and stop when it turns down,
 901   even if the reward is still climbing.
 902
 903**In code:** `true_quality` is the goal, `fit_reward_model` and `proxy_reward` are the flawed reward, `optimise_lengths` trains against either with an optional leash, and `kl_regularised_optimum` with `leash_sweep` gives the closed-form best policy for each β.
 904
 905## In 20 seconds
 906
 907- **RL** learns from a score, not an answer: sample an action from the
 908  policy, get a reward, make rewarded actions more likely.
 909- **REINFORCE**: step along reward × ∇ log π(action); for softmax, ∇ log π
 910  is one-hot minus the probabilities. Unbiased but noisy.
 911- **Baselines** subtract the typical reward, turning rewards into
 912  advantages. Same average gradient, far less variance.
 913- **PPO** reuses each batch for several steps, clips the probability ratio
 914  to 1 ± ε so no step goes too far, and uses a value network as baseline
 915  plus a KL leash to a reference model.
 916- **GRPO** drops the value network: sample a group of answers per prompt
 917  and normalise rewards within the group. With a verifier as reward, it is
 918  how reasoning models are trained.
 919- **Reward hacking**: the policy optimises the reward you wrote, not the
 920  goal you meant. Defend with verifiable rewards, better reward models, a KL
 921  leash and a held-out check of the real goal.
 922
 923## Self-test questions
 924
 925**How does reinforcement learning differ from supervised learning?**
 926Supervised learning is given the correct output for every input and
 927learns to copy it. Reinforcement learning is given only a score for the
 928output it produced, so it must try things, see how they score and shift
 929probability towards what scored well. It fits tasks where judging an
 930answer is easy but writing the perfect one is not.
 931
 932**What is the log-probability trick, and why is it needed?**
 933The gradient of expected reward, Σ ∇π(a) R(a), can't be computed without
 934knowing every action's reward. Rewriting ∇π = π ∇log π turns it into an
 935average over actions sampled from the policy, E[R ∇log π(a)], so each
 936sampled action and its reward give an unbiased estimate of the gradient.
 937
 938**Why does subtracting a baseline not change the expected gradient?**
 939Because E[b ∇log π(a)] = b ∇ Σ π(a) = b ∇ 1 = 0: probabilities always add to
 9401, so the baseline's push averages to nothing. It only removes the noise
 941that comes from rewards being large or all the same sign.
 942
 943**In PPO, what is the probability ratio, and what does clipping it do?**
 944The ratio is the current policy's probability of a sampled action divided
 945by the probability under the policy that sampled it; it measures how far
 946the policy has moved on that sample. Clipping stops counting movement
 947beyond 1 ± ε in the direction the advantage favours, so reusing a batch for
 948several steps can't push the policy far from where the data came from.
 949
 950**Why does PPO's objective take the minimum of the clipped and unclipped terms?**
 951To stay pessimistic. Gains are capped once the ratio leaves the band, but
 952if a step made a bad action more likely, the full penalty still applies, so
 953the objective never rewards a harmful move.
 954
 955**How does GRPO get a baseline without a value network?**
 956It samples several answers to the same prompt and uses their mean reward
 957as the baseline, dividing by their standard deviation to set the scale.
 958Each answer is judged against its siblings, which saves training and
 959serving a second model the size of the policy.
 960
 961**What happens in GRPO when every answer in a group gets the same reward?**
 962Every advantage is zero, so that prompt contributes no gradient. Prompts
 963that are always solved or never solved teach nothing; learning comes from
 964prompts at the edge of the model's ability.
 965
 966**What is reward hacking, and why does a KL penalty help against it?**
 967Reward hacking is the policy maximising the reward as written while the
 968real goal gets worse, usually by finding inputs where a learned reward
 969model is wrong. The KL penalty charges the policy for drifting from the
 970reference model, which keeps it near the kind of outputs the reward model
 971was trained on, where the reward is still trustworthy.
 972
 973**Why are verifiable rewards attractive for training reasoning?**
 974A program that checks the final answer (or runs the tests) has no learned
 975blind spots to exploit and costs nothing to label, so RL can run for a long
 976time against it without the reward drifting away from correctness.
 977
 978## The papers behind this lesson
 979
 980- **Williams, *Simple statistical gradient-following algorithms for
 981  connectionist reinforcement learning* (Machine Learning, 1992)**:
 982  https://link.springer.com/article/10.1007/BF00992696. Introduced
 983  REINFORCE, the log-probability policy gradient with a baseline.
 984  [Annotated companion](../../papers/reinforce.html)
 985- **Schulman et al., *Proximal Policy Optimization Algorithms* (2017)**:
 986  https://arxiv.org/abs/1707.06347. Introduced the clipped probability-ratio
 987  objective that lets each batch be reused for several safe steps.
 988  [Annotated companion](../../papers/ppo.html)
 989- **Ouyang et al., *Training language models to follow instructions with
 990  human feedback* (InstructGPT, 2022)**: https://arxiv.org/abs/2203.02155.
 991  Used PPO with a per-token KL penalty to a reference model to tune a
 992  language model against a learned reward model.
 993  [Annotated companion](../../papers/instructgpt.html)
 994- **Shao et al., *DeepSeekMath: Pushing the Limits of Mathematical
 995  Reasoning in Open Language Models* (2024)**:
 996  https://arxiv.org/abs/2402.03300. Introduced GRPO, replacing PPO's value
 997  network with group-relative advantages.
 998  [Annotated companion](../../papers/deepseekmath-grpo.html)
 999- **DeepSeek-AI, *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs
1000  via Reinforcement Learning* (2025)**: https://arxiv.org/abs/2501.12948.
1001  Showed GRPO with rule-based, verifiable rewards alone can teach a model to
1002  produce long, self-checking chains of thought.
1003  [Annotated companion](../../papers/deepseek-r1.html)
1004- **Gao, Schulman & Hilton, *Scaling Laws for Reward Model
1005  Overoptimization* (2022)**: https://arxiv.org/abs/2210.10760. Measured how
1006  true quality rises and then falls as a policy is optimised further against
1007  a learned reward model.
1008  [Annotated companion](../../papers/reward-model-overoptimization.html)
1009
1010## Further reading
1011
1012- Sutton & Barto, *Reinforcement Learning: An Introduction* (2nd edition, free online): http://incompleteideas.net/book/the-book-2nd.html
1013- OpenAI, *Spinning Up in Deep RL*: https://spinningup.openai.com/
1014- Andrej Karpathy, *Deep Reinforcement Learning: Pong from Pixels*: http://karpathy.github.io/2016/05/31/rl/
1015- Lilian Weng, *Policy Gradient Algorithms*: https://lilianweng.github.io/posts/2018-04-08-policy-gradient/
1016- Hugging Face TRL, *GRPO Trainer*: https://huggingface.co/docs/trl/grpo_trainer
1017- Amodei et al., *Concrete Problems in AI Safety* (2016), section on reward hacking: https://arxiv.org/abs/1606.06565
1018- Schulman et al., *Proximal Policy Optimization Algorithms* (2017): https://arxiv.org/abs/1707.06347
1019- Shao et al., *DeepSeekMath* (2024), which introduces GRPO: https://arxiv.org/abs/2402.03300
1020"""
1021
1022from __future__ import annotations
1023
1024import numpy as np
1025
1026from primer._show import banner, say, table, takeaway
1027from primer.ml.attention import softmax
1028from primer.ml.training_stages import kl_divergence
1029
1030# ---------------------------------------------------------------------------
1031# 1. The setting: a bandit, a policy, and the reward it expects
1032# ---------------------------------------------------------------------------
1033
1034ARMS = ("A", "B", "C")
1035WIN_CHANCES = (0.2, 0.5, 0.8)  # chance each arm pays out 1; C is the one to find
1036
1037
1038class Bandit:
1039    """A row of slot machines. Pulling arm k pays `offset + 1` with chance
1040    `win_chances[k]`, else `offset`.
1041
1042    The agent never sees `win_chances`: it only sees what each pull pays.
1043    That hidden-ness is the whole difficulty of reinforcement learning.
1044    """
1045
1046    def __init__(self, win_chances, offset: float = 0.0, seed: int = 0):
1047        self.win_chances = np.asarray(win_chances, dtype=float)
1048        self.offset = float(offset)
1049        self.rng = np.random.default_rng(seed)
1050
1051    @property
1052    def n_arms(self) -> int:
1053        return len(self.win_chances)
1054
1055    def pull(self, arm: int) -> float:
1056        """Play one arm and return its reward: a noisy, delayed-free score."""
1057        return self.offset + float(self.rng.random() < self.win_chances[arm])
1058
1059
1060def expected_reward(probs: np.ndarray, arm_rewards: np.ndarray) -> float:
1061    """J = Σ_a π(a)·R(a): the average payout of a policy, if you knew every arm's average."""
1062    return float(np.dot(probs, arm_rewards))
1063
1064
1065# ---------------------------------------------------------------------------
1066# 2. Policy gradients: REINFORCE
1067# ---------------------------------------------------------------------------
1068
1069
1070def grad_log_prob(logits: np.ndarray, action: int) -> np.ndarray:
1071    """∇_z log softmax(z)[action] = one_hot(action) − softmax(z).
1072
1073    Raise the chosen action's logit, lower every logit in proportion to how
1074    likely it already was. Shape: same as `logits`.
1075    """
1076    g = -softmax(logits)
1077    g[action] += 1.0
1078    return g
1079
1080
1081def reinforce_step(logits: np.ndarray, action: int, reward: float, lr: float, baseline: float = 0.0) -> np.ndarray:
1082    """One REINFORCE update: z ← z + α·(R − b)·∇ log π(action)."""
1083    return logits + lr * (reward - baseline) * grad_log_prob(logits, action)
1084
1085
1086def worked_reinforce_step() -> np.ndarray:
1087    """The lesson's worked example: a uniform policy over three arms picks C,
1088    earns reward 1, and takes one step at learning rate 0.5.
1089
1090    Logits (0, 0, 0) → (−1/6, −1/6, 1/3); probabilities (1/3 each) → (0.274, 0.274, 0.452).
1091    """
1092    return softmax(reinforce_step(np.zeros(3), action=2, reward=1.0, lr=0.5))
1093
1094
1095def gradient_estimate_stats(logits: np.ndarray, arm_rewards: np.ndarray, baseline: float = 0.0) -> dict[str, np.ndarray]:
1096    """The exact mean and variance of the one-sample estimate (R(a) − b)·∇log π(a).
1097
1098    Rewards here are fixed per arm, so we can enumerate every action instead
1099    of sampling: each action's estimate, weighted by how often the policy
1100    picks it. `mean` is the true gradient of expected reward; `variance` is
1101    how much a single sample swings around it, per logit.
1102    """
1103    probs = softmax(logits)
1104    # One row per action a: (R(a) − b)·∇log π(a). Shape (n_actions, n_logits).
1105    estimates = np.stack([(r - baseline) * grad_log_prob(logits, a) for a, r in enumerate(arm_rewards)])
1106    mean = probs @ estimates
1107    variance = probs @ (estimates - mean) ** 2
1108    return {"mean": mean, "variance": variance}
1109
1110
1111def sampled_gradient_variance(bandit: Bandit, n: int = 4000, baseline: float = 0.0, seed: int = 0) -> float:
1112    """Total variance (summed over logits) of n one-sample gradient estimates,
1113    taken at the uniform starting policy with real noisy pulls."""
1114    # The agent's dice get a stream of their own ([seed, 1]), so they never mirror the bandit's.
1115    rng = np.random.default_rng([seed, 1])
1116    logits = np.zeros(bandit.n_arms)
1117    estimates = []
1118    for _ in range(n):
1119        a = int(rng.choice(bandit.n_arms, p=softmax(logits)))
1120        estimates.append((bandit.pull(a) - baseline) * grad_log_prob(logits, a))
1121    return float(np.var(np.array(estimates), axis=0).sum())
1122
1123
1124def train_reinforce(bandit: Bandit, steps: int = 500, lr: float = 0.1, baseline: bool = False, seed: int = 0) -> dict:
1125    """REINFORCE on a bandit, one pull per step.
1126
1127    With `baseline=True`, each reward is compared with the average of every
1128    reward seen before it (on the first pull there is no history, so the
1129    reward is its own baseline and the step is zero).
1130
1131    Returns `probs` (steps + 1, n_arms), the policy after each step, and
1132    `rewards` (steps,), what each pull paid.
1133    """
1134    # The agent's dice get a stream of their own ([seed, 1]), so they never mirror the bandit's.
1135    rng = np.random.default_rng([seed, 1])
1136    logits = np.zeros(bandit.n_arms)
1137    probs, rewards = [softmax(logits)], []
1138    total = 0.0
1139    for t in range(steps):
1140        a = int(rng.choice(bandit.n_arms, p=softmax(logits)))
1141        r = bandit.pull(a)
1142        b = (total / t if t else r) if baseline else 0.0
1143        logits = reinforce_step(logits, a, r, lr, baseline=b)
1144        total += r
1145        rewards.append(r)
1146        probs.append(softmax(logits))
1147    return {"probs": np.array(probs), "rewards": np.array(rewards)}
1148
1149
1150# ---------------------------------------------------------------------------
1151# 3. PPO: the probability ratio, the clip, and the KL leash
1152# ---------------------------------------------------------------------------
1153
1154
1155def ppo_clipped_objective(ratio, advantage, eps: float = 0.2):
1156    """min(ratio·A, clip(ratio, 1 − ε, 1 + ε)·A), elementwise.
1157
1158    The min takes the more pessimistic of the two: gains are capped once the
1159    ratio leaves the band, penalties never are.
1160    """
1161    ratio, advantage = np.asarray(ratio, dtype=float), np.asarray(advantage, dtype=float)
1162    out = np.minimum(ratio * advantage, np.clip(ratio, 1 - eps, 1 + eps) * advantage)
1163    return float(out) if out.ndim == 0 else out
1164
1165
1166def clipped_policy_update(
1167    logits: np.ndarray,
1168    actions: np.ndarray,
1169    advantages: np.ndarray,
1170    epochs: int = 1,
1171    lr: float = 0.3,
1172    eps: float | None = 0.2,
1173    ref_probs: np.ndarray | None = None,
1174    beta: float = 0.0,
1175) -> tuple[np.ndarray, np.ndarray]:
1176    """Several passes of gradient ascent on the PPO objective over one batch.
1177
1178    The batch (actions, advantages) was sampled from the policy as it was on
1179    entry, the "old" policy. Each pass recomputes every sample's ratio
1180    π_new(a)/π_old(a) and follows the gradient of the clipped objective,
1181    averaged over the batch. `eps=None` switches the clip off. With
1182    `ref_probs` and `beta`, the gradient of −β·KL(π ‖ π_ref) is added too.
1183
1184    Returns (new logits, ratios of shape (epochs, batch) measured at the
1185    start of each pass).
1186    """
1187    logits = logits.astype(float).copy()
1188    old_logp = np.log(softmax(logits))[actions]  # frozen: what the sampler believed
1189    history = []
1190    for _ in range(epochs):
1191        probs = softmax(logits)
1192        ratios = np.exp(np.log(probs)[actions] - old_logp)
1193        history.append(ratios)
1194        grad = np.zeros_like(logits)
1195        for a, adv, ratio in zip(actions, advantages, ratios):
1196            # Outside the band on the side the advantage pushes towards, the clipped
1197            # term is the min and it is flat: no gradient, no further push.
1198            clipped = eps is not None and ((adv > 0 and ratio > 1 + eps) or (adv < 0 and ratio < 1 - eps))
1199            if not clipped:
1200                # d(ratio·A)/dz = A·ratio·∇log π(a), because d ratio = ratio·d log π.
1201                grad += adv * ratio * grad_log_prob(logits, a)
1202        grad /= len(actions)
1203        if ref_probs is not None and beta:
1204            # ∇_z KL(π ‖ π_ref) = π ⊙ (log(π/π_ref) − KL): push back towards the reference.
1205            log_ratio = np.log(probs / ref_probs)
1206            grad -= beta * probs * (log_ratio - probs @ log_ratio)
1207        logits = logits + lr * grad
1208    return logits, np.array(history)
1209
1210
1211def ppo_drift_experiment(epochs: int = 50, lr: float = 0.3, batch: int = 16, seed: int = 1) -> dict[str, np.ndarray]:
1212    """Reuse one batch of 16 pulls for 50 passes, with and without the clip.
1213
1214    Advantages are rewards minus the batch average. Returns the ratios
1215    (epochs, batch) for each run, and each run's final probabilities.
1216    """
1217    rng = np.random.default_rng([seed, 1])  # the agent's own dice, separate from the bandit's
1218    bandit = Bandit(WIN_CHANCES, seed=seed)
1219    actions = rng.choice(3, size=batch)
1220    rewards = np.array([bandit.pull(int(a)) for a in actions])
1221    advantages = rewards - rewards.mean()
1222    out: dict[str, np.ndarray] = {"actions": actions, "rewards": rewards}
1223    for name, eps in (("clipped", 0.2), ("unclipped", None)):
1224        logits, ratios = clipped_policy_update(np.zeros(3), actions, advantages, epochs=epochs, lr=lr, eps=eps)
1225        out[name] = ratios
1226        out[name + "_probs"] = softmax(logits)
1227    return out
1228
1229
1230def kl_penalised_reward(reward: float, logp: float, logp_ref: float, beta: float) -> float:
1231    """R − β·(log π(a) − log π_ref(a)): the per-sample reward RLHF actually optimises."""
1232    return reward - beta * (logp - logp_ref)
1233
1234
1235def kl_regularised_optimum(ref_probs: np.ndarray, rewards: np.ndarray, beta: float) -> np.ndarray:
1236    """The policy that maximises E[R] − β·KL(π ‖ π_ref): π*(a) ∝ π_ref(a)·e^(R(a)/β).
1237
1238    Computed in log space so a tiny β (a huge R/β) cannot overflow.
1239    """
1240    log_w = np.log(ref_probs) + np.asarray(rewards, dtype=float) / beta
1241    return softmax(log_w)
1242
1243
1244# ---------------------------------------------------------------------------
1245# 4. GRPO: group-relative advantages and a verifiable reward
1246# ---------------------------------------------------------------------------
1247
1248# Toy "language model" task: one-token answers (the digits 0 to 9) to additions.
1249ADDITION_PROMPTS = ((1, 2), (2, 2), (3, 1), (4, 3), (2, 5), (3, 3), (1, 7), (4, 5))
1250DIGITS = np.arange(10)
1251
1252
1253def group_advantages(rewards: np.ndarray, eps: float = 1e-8) -> np.ndarray:
1254    """A_i = (r_i − mean(r)) / std(r), within one group of answers to the same prompt.
1255
1256    `eps` keeps an all-equal group (std 0) at exactly zero advantage instead
1257    of dividing by zero. Population std, so hand calculations are exact.
1258    """
1259    rewards = np.asarray(rewards, dtype=float)
1260    return (rewards - rewards.mean()) / (rewards.std() + eps)
1261
1262
1263def verify(prompt: tuple[int, int], answer: int) -> float:
1264    """A verifiable reward: 1 if the answer is the correct sum, else 0. No model, no opinion."""
1265    a, b = prompt
1266    return 1.0 if answer == a + b else 0.0
1267
1268
1269def pretrained_logits(seed: int = 0) -> np.ndarray:
1270    """A weak starting model: (n_prompts, 10) logits over the digits.
1271
1272    It leans a little towards the right answer (+1.0) and its neighbours
1273    (+0.5, the classic off-by-one), with some noise. About 23% accurate.
1274    """
1275    rng = np.random.default_rng(seed)
1276    logits = rng.normal(0.0, 0.3, (len(ADDITION_PROMPTS), len(DIGITS)))
1277    for i, (a, b) in enumerate(ADDITION_PROMPTS):
1278        logits[i, a + b] += 1.0
1279        for near in (a + b - 1, a + b + 1):
1280            if 0 <= near <= 9:
1281                logits[i, near] += 0.5
1282    return logits
1283
1284
1285def _accuracy(logits: np.ndarray) -> float:
1286    probs = softmax(logits)
1287    return float(np.mean([probs[i, a + b] for i, (a, b) in enumerate(ADDITION_PROMPTS)]))
1288
1289
1290def train_grpo(steps: int = 60, group_size: int = 8, lr: float = 0.5, beta: float = 0.0, epochs: int = 1, seed: int = 0) -> dict:
1291    """GRPO on the addition prompts.
1292
1293    Each step, for every prompt: sample `group_size` answers, score them
1294    with `verify`, turn the scores into group-relative advantages, and apply
1295    the clipped update (with a KL leash to the starting model when beta > 0).
1296    No value network anywhere.
1297
1298    Returns `accuracy` (steps + 1,), the average chance of a correct answer,
1299    and `no_signal` (steps,), the share of prompts whose group was all right
1300    or all wrong and so taught nothing.
1301    """
1302    rng = np.random.default_rng(seed)
1303    logits = pretrained_logits()
1304    ref = softmax(logits.copy())
1305    accuracy, no_signal = [_accuracy(logits)], []
1306    for _ in range(steps):
1307        silent = 0
1308        for i, prompt in enumerate(ADDITION_PROMPTS):
1309            answers = rng.choice(DIGITS, size=group_size, p=softmax(logits[i]))
1310            rewards = np.array([verify(prompt, int(a)) for a in answers])
1311            silent += rewards.std() == 0
1312            logits[i], _ = clipped_policy_update(
1313                logits[i], answers, group_advantages(rewards), epochs=epochs, lr=lr, ref_probs=ref[i], beta=beta
1314            )
1315        accuracy.append(_accuracy(logits))
1316        no_signal.append(silent / len(ADDITION_PROMPTS))
1317    return {"accuracy": np.array(accuracy), "no_signal": np.array(no_signal)}
1318
1319
1320# ---------------------------------------------------------------------------
1321# 5. Reward hacking: a flawed reward model for answer length
1322# ---------------------------------------------------------------------------
1323
1324SENTENCES = np.arange(1, 11)  # the "action": how many sentences the answer runs to
1325RATED_LENGTHS = (1, 2, 3, 4)  # the only lengths people rated when the reward model was fit
1326# The starting (fine-tuned) model writes 2 or 3 sentences: a little too short.
1327REFERENCE_LOGITS = -((SENTENCES - 2.5) ** 2) / 6
1328
1329
1330def true_quality(n):
1331    """What we actually want: best at 4 sentences, worse either side, below zero past 8."""
1332    return 1 - ((np.asarray(n, dtype=float) - 4) / 4) ** 2
1333
1334
1335def fit_reward_model(lengths=RATED_LENGTHS) -> tuple[float, float]:
1336    """Least-squares line through the true quality at the rated lengths.
1337
1338    Returns (intercept, slope). On lengths 1 to 4 quality really does rise
1339    with length, so the line says "longer is better", everywhere.
1340    """
1341    x = np.asarray(lengths, dtype=float)
1342    y = true_quality(x)
1343    slope = np.sum((x - x.mean()) * (y - y.mean())) / np.sum((x - x.mean()) ** 2)
1344    return float(y.mean() - slope * x.mean()), float(slope)
1345
1346
1347def proxy_reward(n):
1348    """The learned reward model's score: a straight line in length, extrapolated far past its data."""
1349    intercept, slope = fit_reward_model()
1350    return intercept + slope * np.asarray(n, dtype=float)
1351
1352
1353def optimise_lengths(reward_fn=proxy_reward, steps: int = 300, lr: float = 2.0, beta: float = 0.0) -> dict:
1354    """Gradient ascent on E[reward] − β·KL(π ‖ π_ref) over answer lengths.
1355
1356    Uses the exact expected gradient π ⊙ (R − J) rather than sampled pulls,
1357    so the curves show what the objective rewards, free of sampling noise.
1358    Returns per-step `proxy` and `true` (the policy's average proxy reward
1359    and true quality), `kl` from the reference, and the final `probs`.
1360    """
1361    rewards = reward_fn(SENTENCES)
1362    ref = softmax(REFERENCE_LOGITS)
1363    logits = REFERENCE_LOGITS.astype(float).copy()
1364    proxy, true, kl = [], [], []
1365    for t in range(steps + 1):
1366        probs = softmax(logits)
1367        proxy.append(expected_reward(probs, proxy_reward(SENTENCES)))
1368        true.append(expected_reward(probs, true_quality(SENTENCES)))
1369        kl.append(kl_divergence(probs, ref))
1370        if t == steps:
1371            break
1372        grad = probs * (rewards - probs @ rewards)
1373        log_ratio = np.log(probs / ref)
1374        grad -= beta * probs * (log_ratio - probs @ log_ratio)
1375        logits = logits + lr * grad
1376    return {"proxy": np.array(proxy), "true": np.array(true), "kl": np.array(kl), "probs": softmax(logits)}
1377
1378
1379def leash_sweep(betas=(0.05, 0.1, 0.2, 0.3, 0.5, 1.0, 2.0, 5.0)) -> list[dict]:
1380    """For each β, the best policy under the flawed reward with a KL leash of strength β.
1381
1382    Each row: beta, the policy's average proxy reward and true quality, and
1383    its KL from the reference. Small β lets the policy run to the reward
1384    model's blind spot; large β pins it to the reference.
1385    """
1386    ref = softmax(REFERENCE_LOGITS)
1387    rows = []
1388    for beta in betas:
1389        best = kl_regularised_optimum(ref, proxy_reward(SENTENCES), beta)
1390        rows.append(
1391            dict(
1392                beta=beta,
1393                proxy=expected_reward(best, proxy_reward(SENTENCES)),
1394                true=expected_reward(best, true_quality(SENTENCES)),
1395                kl=kl_divergence(best, ref),
1396            )
1397        )
1398    return rows
1399
1400
1401# ---------------------------------------------------------------------------
1402# 6. Figures (rendered into the HTML docs by `make figures`)
1403# ---------------------------------------------------------------------------
1404
1405
1406def figures() -> dict:
1407    """Plot this lesson's data. matplotlib is imported here, and only here,
1408    so the lesson itself needs nothing beyond NumPy."""
1409    import matplotlib
1410
1411    matplotlib.use("Agg")
1412    import matplotlib.pyplot as plt
1413
1414    BLUE, RED, GREEN, ORANGE, MUTED = "#2563eb", "#dc2626", "#059669", "#d97706", "#9ca3af"
1415    figs = {}
1416
1417    # --- 1. REINFORCE on the bandit -----------------------------------------
1418    history = train_reinforce(Bandit(WIN_CHANCES, seed=0), steps=500, lr=0.1)
1419    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.4))
1420    for k, (name, color) in enumerate(zip(ARMS, (MUTED, ORANGE, BLUE))):
1421        a1.plot(history["probs"][:, k], color=color, label=f"arm {name} (wins {WIN_CHANCES[k]:.0%})")
1422    a1.set(xlabel="pull", ylabel="probability of picking the arm", title="The policy finds the best arm", ylim=(0, 1))
1423    a1.legend(frameon=False)
1424    window = 50
1425    moving = np.convolve(history["rewards"], np.ones(window) / window, mode="valid")
1426    a2.plot(np.arange(window, len(history["rewards"]) + 1), moving, color=BLUE)
1427    a2.axhline(0.8, color=GREEN, ls="--", label="best possible (always C)")
1428    a2.axhline(0.5, color=MUTED, ls=":", label="random pulling")
1429    a2.set(xlabel="pull", ylabel=f"reward, average of last {window}", title="Reward rises as it learns", ylim=(0.3, 0.9))
1430    a2.legend(frameon=False, loc="lower right")
1431    fig.tight_layout()
1432    figs["bandit_learning"] = fig
1433
1434    # --- 2. Baselines: variance, and what it does to learning ----------------
1435    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.4), gridspec_kw={"width_ratios": [1, 1.4]})
1436    variances = [
1437        sampled_gradient_variance(Bandit(WIN_CHANCES, offset=5.0, seed=0), baseline=0.0),
1438        sampled_gradient_variance(Bandit(WIN_CHANCES, offset=5.0, seed=0), baseline=5.5),
1439    ]
1440    a1.bar(["no baseline", "baseline 5.5"], variances, color=[RED, BLUE])
1441    for x, v in enumerate(variances):
1442        a1.text(x, v * 1.15, f"{v:.2f}", ha="center")
1443    a1.set_yscale("log")
1444    a1.set(ylabel="variance of one-pull estimate", title="Rewards of 5 or 6: gradient noise", ylim=(0.05, 60))
1445    rng = np.random.default_rng(0)
1446    for row, (baseline, color, label) in enumerate(((False, RED, "no baseline"), (True, BLUE, "running-average baseline"))):
1447        finals = [train_reinforce(Bandit(WIN_CHANCES, offset=5.0, seed=s), 400, 0.1, baseline=baseline, seed=s)["probs"][-1][2] for s in range(20)]
1448        a2.scatter(finals, row + rng.uniform(-0.15, 0.15, len(finals)), color=color, alpha=0.8)
1449    a2.set_yticks([0, 1], ["no baseline", "baseline"])
1450    a2.set(xlabel="final probability of the best arm, C", title="20 runs each, 400 pulls", xlim=(-0.05, 1.05), ylim=(-0.6, 1.6))
1451    fig.tight_layout()
1452    figs["baselines"] = fig
1453
1454    # --- 3. PPO: the clipped objective and ratio drift -----------------------
1455    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.4))
1456    ratios = np.linspace(0.4, 1.8, 300)
1457    for adv, color in ((1.0, BLUE), (-1.0, RED)):
1458        a1.plot(ratios, ratios * adv, color=color, ls="--", alpha=0.5)
1459        a1.plot(ratios, ppo_clipped_objective(ratios, adv), color=color, label=f"advantage {adv:+.0f}")
1460    a1.axvspan(0.8, 1.2, color=MUTED, alpha=0.2, label="band 1 ± 0.2")
1461    a1.set(xlabel="probability ratio  π_new / π_old", ylabel="objective for one sample", title="The clip: flat outside the band")
1462    a1.legend(frameon=False, loc="upper left")
1463    drift = ppo_drift_experiment()
1464    a2.plot(drift["unclipped"].max(axis=1), color=RED, label="no clip")
1465    a2.plot(drift["clipped"].max(axis=1), color=BLUE, label="clip ε = 0.2")
1466    a2.axhline(1.2, color=MUTED, ls="--")
1467    a2.set(xlabel="pass over the same 16 pulls", ylabel="largest ratio in the batch", title="Reusing one batch 50 times")
1468    a2.legend(frameon=False)
1469    fig.tight_layout()
1470    figs["ppo_clip"] = fig
1471
1472    # --- 4. GRPO on the addition prompts ------------------------------------
1473    run = train_grpo(steps=60, seed=0)
1474    fig, ax = plt.subplots(figsize=(6.5, 3.6))
1475    ax.bar(np.arange(1, 61), run["no_signal"], color=MUTED, alpha=0.6, label="prompts whose group all agreed (no signal)")
1476    ax.plot(np.arange(61), run["accuracy"], color=BLUE, lw=2, label="chance of a correct answer")
1477    ax.set(xlabel="GRPO step (8 answers per prompt)", ylabel="share", title="GRPO with a verifier: 8 addition prompts", ylim=(0, 1.35))
1478    ax.set_yticks(np.linspace(0, 1, 6))
1479    ax.legend(frameon=False, loc="upper left")
1480    figs["grpo"] = fig
1481
1482    # --- 5. The flawed reward model -----------------------------------------
1483    fig, ax = plt.subplots(figsize=(6.5, 3.6))
1484    fine = np.linspace(1, 10, 200)
1485    ax.axvspan(min(RATED_LENGTHS), max(RATED_LENGTHS), color=MUTED, alpha=0.2, label="lengths people rated")
1486    ax.plot(fine, true_quality(fine), color=BLUE, label="true quality")
1487    ax.plot(fine, proxy_reward(fine), color=RED, label="reward model (a fitted line)")
1488    ax.plot(RATED_LENGTHS, true_quality(np.array(RATED_LENGTHS)), "o", color=BLUE)
1489    ax.axhline(0, color="#4b5563", lw=0.8)
1490    ax.set(xlabel="answer length (sentences)", ylabel="score", title="The reward model extrapolates; the truth turns over")
1491    ax.legend(frameon=False, loc="lower left")
1492    figs["length_rewards"] = fig
1493
1494    # --- 6. Reward hacking during training, and the leash sweep --------------
1495    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.6))
1496    for label, reward_fn, beta, color in (
1497        ("flawed reward, no leash", proxy_reward, 0.0, RED),
1498        ("flawed reward, KL leash β = 0.3", proxy_reward, 0.3, ORANGE),
1499        ("verifiable reward (the true goal)", true_quality, 0.0, BLUE),
1500    ):
1501        a1.plot(optimise_lengths(reward_fn, beta=beta)["true"], color=color, label=label)
1502    a1.set(xlabel="training step", ylabel="true quality of the policy", title="Rise, then fall: reward hacking", ylim=(0.6, 1.02))
1503    a1.legend(frameon=False, loc="lower left", fontsize=8)
1504    rows = leash_sweep(betas=np.geomspace(0.04, 20, 40))
1505    kls = [r["kl"] for r in rows]
1506    a2.plot(kls, [r["proxy"] for r in rows], color=RED, label="reward model's score")
1507    a2.plot(kls, [r["true"] for r in rows], color=BLUE, label="true quality")
1508    a2.set_xscale("log")
1509    a2.set(xlabel="KL from the reference (looser leash →)", ylabel="average score", title="Best policy for each leash strength")
1510    a2.legend(frameon=False, loc="lower left")
1511    fig.tight_layout()
1512    figs["reward_hacking"] = fig
1513
1514    return figs
1515
1516
1517# ---------------------------------------------------------------------------
1518# 7. Narrated walkthrough
1519# ---------------------------------------------------------------------------
1520
1521
1522def demo() -> None:
1523    banner("1. A bandit: three slot machines, one hidden best")
1524    say(
1525        """
1526        Arms A, B and C pay 1 with chance 20%, 50% and 80%, else 0. The agent
1527        is not told this. A policy that picks uniformly earns 0.5 per pull on
1528        average; always picking C earns 0.8.
1529        """
1530    )
1531    table(
1532        ["policy", "expected reward J"],
1533        [("uniform", expected_reward(np.full(3, 1 / 3), np.array(WIN_CHANCES))), ("always C", expected_reward(np.array([0, 0, 1.0]), np.array(WIN_CHANCES)))],
1534        floatfmt=".2f",
1535    )
1536
1537    banner("2. REINFORCE: one step, by hand")
1538    say("Uniform logits (0, 0, 0). The agent pulls C and wins: reward 1. Step = 0.5 × 1 × (one-hot − π).")
1539    table(["arm", "∇ log π(C)", "new probability"], zip(ARMS, grad_log_prob(np.zeros(3), 2), worked_reinforce_step()), floatfmt=".3f")
1540    history = train_reinforce(Bandit(WIN_CHANCES, seed=0), steps=500, lr=0.1)
1541    say(
1542        f"""
1543        Now 500 pulls. Reward over the first 100: {history['rewards'][:100].mean():.2f};
1544        over the last 100: {history['rewards'][-100:].mean():.2f}. Final policy:
1545        """
1546    )
1547    table(["arm", "win chance", "final probability"], zip(ARMS, WIN_CHANCES, history["probs"][-1]), floatfmt=".3f")
1548    takeaway("Step along reward × ∇ log π(action): rewarded actions become more likely, and on average the best arm wins.")
1549
1550    banner("3. Baselines: same average gradient, far less noise")
1551    for b in (0.0, 11.0):
1552        stats = gradient_estimate_stats(np.zeros(2), np.array([10.0, 12.0]), baseline=b)
1553        say(f"Rewards 10 and 12, baseline {b:g}: mean gradient for arm 2 = {stats['mean'][1]:+.2f}, variance = {stats['variance'][1]:.2f}")
1554    wrong = {
1555        bl: np.mean([train_reinforce(Bandit(WIN_CHANCES, offset=5.0, seed=s), 400, 0.1, baseline=bl, seed=s)["probs"][-1][2] < 0.5 for s in range(20)])
1556        for bl in (False, True)
1557    }
1558    say(
1559        f"""
1560        Offset every reward by 5. Of 20 training runs, {wrong[False]:.0%} lock onto
1561        a worse arm without a baseline, and {wrong[True]:.0%} with a running-average baseline.
1562        """
1563    )
1564    takeaway("Subtract the typical reward: the advantage says 'better or worse than usual', which is the only thing that matters.")
1565
1566    banner("4. PPO: the probability ratio and the clip")
1567    table(
1568        ["ratio", "advantage", "ratio × A", "clipped objective"],
1569        [(r, a, r * a, ppo_clipped_objective(r, a)) for r, a in ((1.1, 3.0), (1.5, 2.0), (0.5, -1.0), (1.5, -1.0))],
1570        floatfmt=".2f",
1571    )
1572    drift = ppo_drift_experiment()
1573    say(
1574        f"""
1575        One batch of 16 pulls, reused for 50 passes. Without the clip the largest
1576        ratio reaches {drift['unclipped'].max():.2f} and C's probability {drift['unclipped_probs'][2]:.2f}
1577        from sixteen pulls. With it, the ratio stops at {drift['clipped'].max():.2f} and C
1578        sits at {drift['clipped_probs'][2]:.2f}.
1579        """
1580    )
1581    say(f"KL leash: reward 1.0 for an answer at probability 0.6 vs reference 0.3, β = 0.1 → {kl_penalised_reward(1.0, np.log(0.6), np.log(0.3), 0.1):.3f}")
1582    takeaway("PPO reuses expensive samples for several steps, and the clip keeps each batch from pulling the policy too far.")
1583
1584    banner("5. GRPO: advantages from a group of answers, no value network")
1585    table(
1586        ["group rewards", "advantages"],
1587        [(str(r), np.array2string(group_advantages(np.array(r, dtype=float)), precision=2)) for r in ([1, 0, 0, 1], [1, 0, 0, 0], [1, 1, 1, 1])],
1588    )
1589    run = train_grpo(steps=60, seed=0)
1590    say(
1591        f"""
1592        GRPO with a verifier on 8 addition prompts, 8 answers each: accuracy goes from
1593        {run['accuracy'][0]:.0%} to {run['accuracy'][-1]:.0%} in 60 steps. By the end,
1594        {run['no_signal'][-10:].mean():.0%} of groups are unanimous and teach nothing.
1595        """
1596    )
1597    takeaway("Compare each answer with its siblings: the group mean is the baseline, and a verifier is the reward.")
1598
1599    banner("6. Reward hacking: a reward model that loves length")
1600    intercept, slope = fit_reward_model()
1601    say(f"Fitted on answers of 1 to 4 sentences, the reward model is {intercept:.4f} + {slope:.4f} × sentences.")
1602    table(
1603        ["sentences", "true quality", "reward model"],
1604        [(int(n), true_quality(n), proxy_reward(n)) for n in (1, 2, 3, 4, 6, 8, 10)],
1605        floatfmt=".2f",
1606    )
1607    free, leashed, true_run = optimise_lengths(), optimise_lengths(beta=0.3), optimise_lengths(true_quality)
1608    table(
1609        ["training against", "true quality: start", "peak", "end"],
1610        [
1611            (name, r["true"][0], r["true"].max(), r["true"][-1])
1612            for name, r in (("flawed reward", free), ("flawed reward + KL β=0.3", leashed), ("verifiable reward", true_run))
1613        ],
1614        floatfmt=".2f",
1615    )
1616    say(
1617        f"""
1618        Against the flawed reward, quality rises while lengthening helps, then falls below
1619        where it started as the policy settles on {SENTENCES[free['probs'].argmax()]}-sentence answers,
1620        even as the reward model's score climbs from {free['proxy'][0]:.2f} to {free['proxy'][-1]:.2f}.
1621        """
1622    )
1623    takeaway(
1624        "The policy optimises the reward you wrote, not the goal you meant. Prefer verifiable rewards, "
1625        "keep a KL leash, and watch a held-out measure of the real goal."
1626    )
1627
1628
1629if __name__ == "__main__":
1630    demo()
Level 3: the code, function by function.
ARMS = ('A', 'B', 'C')
WIN_CHANCES = (0.2, 0.5, 0.8)
class Bandit: on GitHub
1039class Bandit:
1040    """A row of slot machines. Pulling arm k pays `offset + 1` with chance
1041    `win_chances[k]`, else `offset`.
1042
1043    The agent never sees `win_chances`: it only sees what each pull pays.
1044    That hidden-ness is the whole difficulty of reinforcement learning.
1045    """
1046
1047    def __init__(self, win_chances, offset: float = 0.0, seed: int = 0):
1048        self.win_chances = np.asarray(win_chances, dtype=float)
1049        self.offset = float(offset)
1050        self.rng = np.random.default_rng(seed)
1051
1052    @property
1053    def n_arms(self) -> int:
1054        return len(self.win_chances)
1055
1056    def pull(self, arm: int) -> float:
1057        """Play one arm and return its reward: a noisy, delayed-free score."""
1058        return self.offset + float(self.rng.random() < self.win_chances[arm])

A row of slot machines. Pulling arm k pays offset + 1 with chance win_chances[k], else offset.

The agent never sees win_chances: it only sees what each pull pays. That hidden-ness is the whole difficulty of reinforcement learning.

Bandit(win_chances, offset: float = 0.0, seed: int = 0) on GitHub
1047    def __init__(self, win_chances, offset: float = 0.0, seed: int = 0):
1048        self.win_chances = np.asarray(win_chances, dtype=float)
1049        self.offset = float(offset)
1050        self.rng = np.random.default_rng(seed)
offset
rng
n_arms: int on GitHub
1052    @property
1053    def n_arms(self) -> int:
1054        return len(self.win_chances)
def pull(self, arm: int) -> float: on GitHub
1056    def pull(self, arm: int) -> float:
1057        """Play one arm and return its reward: a noisy, delayed-free score."""
1058        return self.offset + float(self.rng.random() < self.win_chances[arm])

Play one arm and return its reward: a noisy, delayed-free score.

def expected_reward(probs: numpy.ndarray, arm_rewards: numpy.ndarray) -> float: on GitHub
1061def expected_reward(probs: np.ndarray, arm_rewards: np.ndarray) -> float:
1062    """J = Σ_a π(a)·R(a): the average payout of a policy, if you knew every arm's average."""
1063    return float(np.dot(probs, arm_rewards))

J = Σ_a π(a)·R(a): the average payout of a policy, if you knew every arm's average.

def grad_log_prob(logits: numpy.ndarray, action: int) -> numpy.ndarray: on GitHub
1071def grad_log_prob(logits: np.ndarray, action: int) -> np.ndarray:
1072    """∇_z log softmax(z)[action] = one_hot(action) − softmax(z).
1073
1074    Raise the chosen action's logit, lower every logit in proportion to how
1075    likely it already was. Shape: same as `logits`.
1076    """
1077    g = -softmax(logits)
1078    g[action] += 1.0
1079    return g

∇_z log softmax(z)[action] = one_hot(action) − softmax(z).

Raise the chosen action's logit, lower every logit in proportion to how likely it already was. Shape: same as logits.

def reinforce_step( logits: numpy.ndarray, action: int, reward: float, lr: float, baseline: float = 0.0) -> numpy.ndarray: on GitHub
1082def reinforce_step(logits: np.ndarray, action: int, reward: float, lr: float, baseline: float = 0.0) -> np.ndarray:
1083    """One REINFORCE update: z ← z + α·(R − b)·∇ log π(action)."""
1084    return logits + lr * (reward - baseline) * grad_log_prob(logits, action)

One REINFORCE update: z ← z + α·(R − b)·∇ log π(action).

def worked_reinforce_step() -> numpy.ndarray: on GitHub
1087def worked_reinforce_step() -> np.ndarray:
1088    """The lesson's worked example: a uniform policy over three arms picks C,
1089    earns reward 1, and takes one step at learning rate 0.5.
1090
1091    Logits (0, 0, 0) → (−1/6, −1/6, 1/3); probabilities (1/3 each) → (0.274, 0.274, 0.452).
1092    """
1093    return softmax(reinforce_step(np.zeros(3), action=2, reward=1.0, lr=0.5))

The lesson's worked example: a uniform policy over three arms picks C, earns reward 1, and takes one step at learning rate 0.5.

Logits (0, 0, 0) → (−1/6, −1/6, 1/3); probabilities (1/3 each) → (0.274, 0.274, 0.452).

def gradient_estimate_stats( logits: numpy.ndarray, arm_rewards: numpy.ndarray, baseline: float = 0.0) -> dict[str, numpy.ndarray]: on GitHub
1096def gradient_estimate_stats(logits: np.ndarray, arm_rewards: np.ndarray, baseline: float = 0.0) -> dict[str, np.ndarray]:
1097    """The exact mean and variance of the one-sample estimate (R(a) − b)·∇log π(a).
1098
1099    Rewards here are fixed per arm, so we can enumerate every action instead
1100    of sampling: each action's estimate, weighted by how often the policy
1101    picks it. `mean` is the true gradient of expected reward; `variance` is
1102    how much a single sample swings around it, per logit.
1103    """
1104    probs = softmax(logits)
1105    # One row per action a: (R(a) − b)·∇log π(a). Shape (n_actions, n_logits).
1106    estimates = np.stack([(r - baseline) * grad_log_prob(logits, a) for a, r in enumerate(arm_rewards)])
1107    mean = probs @ estimates
1108    variance = probs @ (estimates - mean) ** 2
1109    return {"mean": mean, "variance": variance}

The exact mean and variance of the one-sample estimate (R(a) − b)·∇log π(a).

Rewards here are fixed per arm, so we can enumerate every action instead of sampling: each action's estimate, weighted by how often the policy picks it. mean is the true gradient of expected reward; variance is how much a single sample swings around it, per logit.

def sampled_gradient_variance( bandit: Bandit, n: int = 4000, baseline: float = 0.0, seed: int = 0) -> float: on GitHub
1112def sampled_gradient_variance(bandit: Bandit, n: int = 4000, baseline: float = 0.0, seed: int = 0) -> float:
1113    """Total variance (summed over logits) of n one-sample gradient estimates,
1114    taken at the uniform starting policy with real noisy pulls."""
1115    # The agent's dice get a stream of their own ([seed, 1]), so they never mirror the bandit's.
1116    rng = np.random.default_rng([seed, 1])
1117    logits = np.zeros(bandit.n_arms)
1118    estimates = []
1119    for _ in range(n):
1120        a = int(rng.choice(bandit.n_arms, p=softmax(logits)))
1121        estimates.append((bandit.pull(a) - baseline) * grad_log_prob(logits, a))
1122    return float(np.var(np.array(estimates), axis=0).sum())

Total variance (summed over logits) of n one-sample gradient estimates, taken at the uniform starting policy with real noisy pulls.

def train_reinforce( bandit: Bandit, steps: int = 500, lr: float = 0.1, baseline: bool = False, seed: int = 0) -> dict: on GitHub
1125def train_reinforce(bandit: Bandit, steps: int = 500, lr: float = 0.1, baseline: bool = False, seed: int = 0) -> dict:
1126    """REINFORCE on a bandit, one pull per step.
1127
1128    With `baseline=True`, each reward is compared with the average of every
1129    reward seen before it (on the first pull there is no history, so the
1130    reward is its own baseline and the step is zero).
1131
1132    Returns `probs` (steps + 1, n_arms), the policy after each step, and
1133    `rewards` (steps,), what each pull paid.
1134    """
1135    # The agent's dice get a stream of their own ([seed, 1]), so they never mirror the bandit's.
1136    rng = np.random.default_rng([seed, 1])
1137    logits = np.zeros(bandit.n_arms)
1138    probs, rewards = [softmax(logits)], []
1139    total = 0.0
1140    for t in range(steps):
1141        a = int(rng.choice(bandit.n_arms, p=softmax(logits)))
1142        r = bandit.pull(a)
1143        b = (total / t if t else r) if baseline else 0.0
1144        logits = reinforce_step(logits, a, r, lr, baseline=b)
1145        total += r
1146        rewards.append(r)
1147        probs.append(softmax(logits))
1148    return {"probs": np.array(probs), "rewards": np.array(rewards)}

REINFORCE on a bandit, one pull per step.

With baseline=True, each reward is compared with the average of every reward seen before it (on the first pull there is no history, so the reward is its own baseline and the step is zero).

Returns probs (steps + 1, n_arms), the policy after each step, and rewards (steps,), what each pull paid.

def ppo_clipped_objective(ratio, advantage, eps: float = 0.2): on GitHub
1156def ppo_clipped_objective(ratio, advantage, eps: float = 0.2):
1157    """min(ratio·A, clip(ratio, 1 − ε, 1 + ε)·A), elementwise.
1158
1159    The min takes the more pessimistic of the two: gains are capped once the
1160    ratio leaves the band, penalties never are.
1161    """
1162    ratio, advantage = np.asarray(ratio, dtype=float), np.asarray(advantage, dtype=float)
1163    out = np.minimum(ratio * advantage, np.clip(ratio, 1 - eps, 1 + eps) * advantage)
1164    return float(out) if out.ndim == 0 else out

min(ratio·A, clip(ratio, 1 − ε, 1 + ε)·A), elementwise.

The min takes the more pessimistic of the two: gains are capped once the ratio leaves the band, penalties never are.

def clipped_policy_update( logits: numpy.ndarray, actions: numpy.ndarray, advantages: numpy.ndarray, epochs: int = 1, lr: float = 0.3, eps: float | None = 0.2, ref_probs: numpy.ndarray | None = None, beta: float = 0.0) -> tuple[numpy.ndarray, numpy.ndarray]: on GitHub
1167def clipped_policy_update(
1168    logits: np.ndarray,
1169    actions: np.ndarray,
1170    advantages: np.ndarray,
1171    epochs: int = 1,
1172    lr: float = 0.3,
1173    eps: float | None = 0.2,
1174    ref_probs: np.ndarray | None = None,
1175    beta: float = 0.0,
1176) -> tuple[np.ndarray, np.ndarray]:
1177    """Several passes of gradient ascent on the PPO objective over one batch.
1178
1179    The batch (actions, advantages) was sampled from the policy as it was on
1180    entry, the "old" policy. Each pass recomputes every sample's ratio
1181    π_new(a)/π_old(a) and follows the gradient of the clipped objective,
1182    averaged over the batch. `eps=None` switches the clip off. With
1183    `ref_probs` and `beta`, the gradient of −β·KL(π ‖ π_ref) is added too.
1184
1185    Returns (new logits, ratios of shape (epochs, batch) measured at the
1186    start of each pass).
1187    """
1188    logits = logits.astype(float).copy()
1189    old_logp = np.log(softmax(logits))[actions]  # frozen: what the sampler believed
1190    history = []
1191    for _ in range(epochs):
1192        probs = softmax(logits)
1193        ratios = np.exp(np.log(probs)[actions] - old_logp)
1194        history.append(ratios)
1195        grad = np.zeros_like(logits)
1196        for a, adv, ratio in zip(actions, advantages, ratios):
1197            # Outside the band on the side the advantage pushes towards, the clipped
1198            # term is the min and it is flat: no gradient, no further push.
1199            clipped = eps is not None and ((adv > 0 and ratio > 1 + eps) or (adv < 0 and ratio < 1 - eps))
1200            if not clipped:
1201                # d(ratio·A)/dz = A·ratio·∇log π(a), because d ratio = ratio·d log π.
1202                grad += adv * ratio * grad_log_prob(logits, a)
1203        grad /= len(actions)
1204        if ref_probs is not None and beta:
1205            # ∇_z KL(π ‖ π_ref) = π ⊙ (log(π/π_ref) − KL): push back towards the reference.
1206            log_ratio = np.log(probs / ref_probs)
1207            grad -= beta * probs * (log_ratio - probs @ log_ratio)
1208        logits = logits + lr * grad
1209    return logits, np.array(history)

Several passes of gradient ascent on the PPO objective over one batch.

The batch (actions, advantages) was sampled from the policy as it was on entry, the "old" policy. Each pass recomputes every sample's ratio π_new(a)/π_old(a) and follows the gradient of the clipped objective, averaged over the batch. eps=None switches the clip off. With ref_probs and beta, the gradient of −β·KL(π ‖ π_ref) is added too.

Returns (new logits, ratios of shape (epochs, batch) measured at the start of each pass).

def ppo_drift_experiment( epochs: int = 50, lr: float = 0.3, batch: int = 16, seed: int = 1) -> dict[str, numpy.ndarray]: on GitHub
1212def ppo_drift_experiment(epochs: int = 50, lr: float = 0.3, batch: int = 16, seed: int = 1) -> dict[str, np.ndarray]:
1213    """Reuse one batch of 16 pulls for 50 passes, with and without the clip.
1214
1215    Advantages are rewards minus the batch average. Returns the ratios
1216    (epochs, batch) for each run, and each run's final probabilities.
1217    """
1218    rng = np.random.default_rng([seed, 1])  # the agent's own dice, separate from the bandit's
1219    bandit = Bandit(WIN_CHANCES, seed=seed)
1220    actions = rng.choice(3, size=batch)
1221    rewards = np.array([bandit.pull(int(a)) for a in actions])
1222    advantages = rewards - rewards.mean()
1223    out: dict[str, np.ndarray] = {"actions": actions, "rewards": rewards}
1224    for name, eps in (("clipped", 0.2), ("unclipped", None)):
1225        logits, ratios = clipped_policy_update(np.zeros(3), actions, advantages, epochs=epochs, lr=lr, eps=eps)
1226        out[name] = ratios
1227        out[name + "_probs"] = softmax(logits)
1228    return out

Reuse one batch of 16 pulls for 50 passes, with and without the clip.

Advantages are rewards minus the batch average. Returns the ratios (epochs, batch) for each run, and each run's final probabilities.

def kl_penalised_reward(reward: float, logp: float, logp_ref: float, beta: float) -> float: on GitHub
1231def kl_penalised_reward(reward: float, logp: float, logp_ref: float, beta: float) -> float:
1232    """R − β·(log π(a) − log π_ref(a)): the per-sample reward RLHF actually optimises."""
1233    return reward - beta * (logp - logp_ref)

R − β·(log π(a) − log π_ref(a)): the per-sample reward RLHF actually optimises.

def kl_regularised_optimum( ref_probs: numpy.ndarray, rewards: numpy.ndarray, beta: float) -> numpy.ndarray: on GitHub
1236def kl_regularised_optimum(ref_probs: np.ndarray, rewards: np.ndarray, beta: float) -> np.ndarray:
1237    """The policy that maximises E[R] − β·KL(π ‖ π_ref): π*(a) ∝ π_ref(a)·e^(R(a)/β).
1238
1239    Computed in log space so a tiny β (a huge R/β) cannot overflow.
1240    """
1241    log_w = np.log(ref_probs) + np.asarray(rewards, dtype=float) / beta
1242    return softmax(log_w)

The policy that maximises E[R] − β·KL(π ‖ π_ref): π*(a) ∝ π_ref(a)·e^(R(a)/β).

Computed in log space so a tiny β (a huge R/β) cannot overflow.

ADDITION_PROMPTS = ((1, 2), (2, 2), (3, 1), (4, 3), (2, 5), (3, 3), (1, 7), (4, 5))
DIGITS = array([0, 1, 2, 3, 4, 5, 6, 7, 8, 9])
def group_advantages(rewards: numpy.ndarray, eps: float = 1e-08) -> numpy.ndarray: on GitHub
1254def group_advantages(rewards: np.ndarray, eps: float = 1e-8) -> np.ndarray:
1255    """A_i = (r_i − mean(r)) / std(r), within one group of answers to the same prompt.
1256
1257    `eps` keeps an all-equal group (std 0) at exactly zero advantage instead
1258    of dividing by zero. Population std, so hand calculations are exact.
1259    """
1260    rewards = np.asarray(rewards, dtype=float)
1261    return (rewards - rewards.mean()) / (rewards.std() + eps)

A_i = (r_i − mean(r)) / std(r), within one group of answers to the same prompt.

eps keeps an all-equal group (std 0) at exactly zero advantage instead of dividing by zero. Population std, so hand calculations are exact.

def verify(prompt: tuple[int, int], answer: int) -> float: on GitHub
1264def verify(prompt: tuple[int, int], answer: int) -> float:
1265    """A verifiable reward: 1 if the answer is the correct sum, else 0. No model, no opinion."""
1266    a, b = prompt
1267    return 1.0 if answer == a + b else 0.0

A verifiable reward: 1 if the answer is the correct sum, else 0. No model, no opinion.

def pretrained_logits(seed: int = 0) -> numpy.ndarray: on GitHub
1270def pretrained_logits(seed: int = 0) -> np.ndarray:
1271    """A weak starting model: (n_prompts, 10) logits over the digits.
1272
1273    It leans a little towards the right answer (+1.0) and its neighbours
1274    (+0.5, the classic off-by-one), with some noise. About 23% accurate.
1275    """
1276    rng = np.random.default_rng(seed)
1277    logits = rng.normal(0.0, 0.3, (len(ADDITION_PROMPTS), len(DIGITS)))
1278    for i, (a, b) in enumerate(ADDITION_PROMPTS):
1279        logits[i, a + b] += 1.0
1280        for near in (a + b - 1, a + b + 1):
1281            if 0 <= near <= 9:
1282                logits[i, near] += 0.5
1283    return logits

A weak starting model: (n_prompts, 10) logits over the digits.

It leans a little towards the right answer (+1.0) and its neighbours (+0.5, the classic off-by-one), with some noise. About 23% accurate.

def train_grpo( steps: int = 60, group_size: int = 8, lr: float = 0.5, beta: float = 0.0, epochs: int = 1, seed: int = 0) -> dict: on GitHub
1291def train_grpo(steps: int = 60, group_size: int = 8, lr: float = 0.5, beta: float = 0.0, epochs: int = 1, seed: int = 0) -> dict:
1292    """GRPO on the addition prompts.
1293
1294    Each step, for every prompt: sample `group_size` answers, score them
1295    with `verify`, turn the scores into group-relative advantages, and apply
1296    the clipped update (with a KL leash to the starting model when beta > 0).
1297    No value network anywhere.
1298
1299    Returns `accuracy` (steps + 1,), the average chance of a correct answer,
1300    and `no_signal` (steps,), the share of prompts whose group was all right
1301    or all wrong and so taught nothing.
1302    """
1303    rng = np.random.default_rng(seed)
1304    logits = pretrained_logits()
1305    ref = softmax(logits.copy())
1306    accuracy, no_signal = [_accuracy(logits)], []
1307    for _ in range(steps):
1308        silent = 0
1309        for i, prompt in enumerate(ADDITION_PROMPTS):
1310            answers = rng.choice(DIGITS, size=group_size, p=softmax(logits[i]))
1311            rewards = np.array([verify(prompt, int(a)) for a in answers])
1312            silent += rewards.std() == 0
1313            logits[i], _ = clipped_policy_update(
1314                logits[i], answers, group_advantages(rewards), epochs=epochs, lr=lr, ref_probs=ref[i], beta=beta
1315            )
1316        accuracy.append(_accuracy(logits))
1317        no_signal.append(silent / len(ADDITION_PROMPTS))
1318    return {"accuracy": np.array(accuracy), "no_signal": np.array(no_signal)}

GRPO on the addition prompts.

Each step, for every prompt: sample group_size answers, score them with verify, turn the scores into group-relative advantages, and apply the clipped update (with a KL leash to the starting model when beta > 0). No value network anywhere.

Returns accuracy (steps + 1,), the average chance of a correct answer, and no_signal (steps,), the share of prompts whose group was all right or all wrong and so taught nothing.

SENTENCES = array([ 1, 2, 3, 4, 5, 6, 7, 8, 9, 10])
RATED_LENGTHS = (1, 2, 3, 4)
REFERENCE_LOGITS = array([-0.375 , -0.04166667, -0.04166667, -0.375 , -1.04166667, -2.04166667, -3.375 , -5.04166667, -7.04166667, -9.375 ])
def true_quality(n): on GitHub
1331def true_quality(n):
1332    """What we actually want: best at 4 sentences, worse either side, below zero past 8."""
1333    return 1 - ((np.asarray(n, dtype=float) - 4) / 4) ** 2

What we actually want: best at 4 sentences, worse either side, below zero past 8.

def fit_reward_model(lengths=(1, 2, 3, 4)) -> tuple[float, float]: on GitHub
1336def fit_reward_model(lengths=RATED_LENGTHS) -> tuple[float, float]:
1337    """Least-squares line through the true quality at the rated lengths.
1338
1339    Returns (intercept, slope). On lengths 1 to 4 quality really does rise
1340    with length, so the line says "longer is better", everywhere.
1341    """
1342    x = np.asarray(lengths, dtype=float)
1343    y = true_quality(x)
1344    slope = np.sum((x - x.mean()) * (y - y.mean())) / np.sum((x - x.mean()) ** 2)
1345    return float(y.mean() - slope * x.mean()), float(slope)

Least-squares line through the true quality at the rated lengths.

Returns (intercept, slope). On lengths 1 to 4 quality really does rise with length, so the line says "longer is better", everywhere.

def proxy_reward(n): on GitHub
1348def proxy_reward(n):
1349    """The learned reward model's score: a straight line in length, extrapolated far past its data."""
1350    intercept, slope = fit_reward_model()
1351    return intercept + slope * np.asarray(n, dtype=float)

The learned reward model's score: a straight line in length, extrapolated far past its data.

def optimise_lengths( reward_fn=<function proxy_reward>, steps: int = 300, lr: float = 2.0, beta: float = 0.0) -> dict: on GitHub
1354def optimise_lengths(reward_fn=proxy_reward, steps: int = 300, lr: float = 2.0, beta: float = 0.0) -> dict:
1355    """Gradient ascent on E[reward] − β·KL(π ‖ π_ref) over answer lengths.
1356
1357    Uses the exact expected gradient π ⊙ (R − J) rather than sampled pulls,
1358    so the curves show what the objective rewards, free of sampling noise.
1359    Returns per-step `proxy` and `true` (the policy's average proxy reward
1360    and true quality), `kl` from the reference, and the final `probs`.
1361    """
1362    rewards = reward_fn(SENTENCES)
1363    ref = softmax(REFERENCE_LOGITS)
1364    logits = REFERENCE_LOGITS.astype(float).copy()
1365    proxy, true, kl = [], [], []
1366    for t in range(steps + 1):
1367        probs = softmax(logits)
1368        proxy.append(expected_reward(probs, proxy_reward(SENTENCES)))
1369        true.append(expected_reward(probs, true_quality(SENTENCES)))
1370        kl.append(kl_divergence(probs, ref))
1371        if t == steps:
1372            break
1373        grad = probs * (rewards - probs @ rewards)
1374        log_ratio = np.log(probs / ref)
1375        grad -= beta * probs * (log_ratio - probs @ log_ratio)
1376        logits = logits + lr * grad
1377    return {"proxy": np.array(proxy), "true": np.array(true), "kl": np.array(kl), "probs": softmax(logits)}

Gradient ascent on E[reward] − β·KL(π ‖ π_ref) over answer lengths.

Uses the exact expected gradient π ⊙ (R − J) rather than sampled pulls, so the curves show what the objective rewards, free of sampling noise. Returns per-step proxy and true (the policy's average proxy reward and true quality), kl from the reference, and the final probs.

def leash_sweep(betas=(0.05, 0.1, 0.2, 0.3, 0.5, 1.0, 2.0, 5.0)) -> list[dict]: on GitHub
1380def leash_sweep(betas=(0.05, 0.1, 0.2, 0.3, 0.5, 1.0, 2.0, 5.0)) -> list[dict]:
1381    """For each β, the best policy under the flawed reward with a KL leash of strength β.
1382
1383    Each row: beta, the policy's average proxy reward and true quality, and
1384    its KL from the reference. Small β lets the policy run to the reward
1385    model's blind spot; large β pins it to the reference.
1386    """
1387    ref = softmax(REFERENCE_LOGITS)
1388    rows = []
1389    for beta in betas:
1390        best = kl_regularised_optimum(ref, proxy_reward(SENTENCES), beta)
1391        rows.append(
1392            dict(
1393                beta=beta,
1394                proxy=expected_reward(best, proxy_reward(SENTENCES)),
1395                true=expected_reward(best, true_quality(SENTENCES)),
1396                kl=kl_divergence(best, ref),
1397            )
1398        )
1399    return rows

For each β, the best policy under the flawed reward with a KL leash of strength β.

Each row: beta, the policy's average proxy reward and true quality, and its KL from the reference. Small β lets the policy run to the reward model's blind spot; large β pins it to the reference.

def figures() -> dict: on GitHub
1407def figures() -> dict:
1408    """Plot this lesson's data. matplotlib is imported here, and only here,
1409    so the lesson itself needs nothing beyond NumPy."""
1410    import matplotlib
1411
1412    matplotlib.use("Agg")
1413    import matplotlib.pyplot as plt
1414
1415    BLUE, RED, GREEN, ORANGE, MUTED = "#2563eb", "#dc2626", "#059669", "#d97706", "#9ca3af"
1416    figs = {}
1417
1418    # --- 1. REINFORCE on the bandit -----------------------------------------
1419    history = train_reinforce(Bandit(WIN_CHANCES, seed=0), steps=500, lr=0.1)
1420    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.4))
1421    for k, (name, color) in enumerate(zip(ARMS, (MUTED, ORANGE, BLUE))):
1422        a1.plot(history["probs"][:, k], color=color, label=f"arm {name} (wins {WIN_CHANCES[k]:.0%})")
1423    a1.set(xlabel="pull", ylabel="probability of picking the arm", title="The policy finds the best arm", ylim=(0, 1))
1424    a1.legend(frameon=False)
1425    window = 50
1426    moving = np.convolve(history["rewards"], np.ones(window) / window, mode="valid")
1427    a2.plot(np.arange(window, len(history["rewards"]) + 1), moving, color=BLUE)
1428    a2.axhline(0.8, color=GREEN, ls="--", label="best possible (always C)")
1429    a2.axhline(0.5, color=MUTED, ls=":", label="random pulling")
1430    a2.set(xlabel="pull", ylabel=f"reward, average of last {window}", title="Reward rises as it learns", ylim=(0.3, 0.9))
1431    a2.legend(frameon=False, loc="lower right")
1432    fig.tight_layout()
1433    figs["bandit_learning"] = fig
1434
1435    # --- 2. Baselines: variance, and what it does to learning ----------------
1436    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.4), gridspec_kw={"width_ratios": [1, 1.4]})
1437    variances = [
1438        sampled_gradient_variance(Bandit(WIN_CHANCES, offset=5.0, seed=0), baseline=0.0),
1439        sampled_gradient_variance(Bandit(WIN_CHANCES, offset=5.0, seed=0), baseline=5.5),
1440    ]
1441    a1.bar(["no baseline", "baseline 5.5"], variances, color=[RED, BLUE])
1442    for x, v in enumerate(variances):
1443        a1.text(x, v * 1.15, f"{v:.2f}", ha="center")
1444    a1.set_yscale("log")
1445    a1.set(ylabel="variance of one-pull estimate", title="Rewards of 5 or 6: gradient noise", ylim=(0.05, 60))
1446    rng = np.random.default_rng(0)
1447    for row, (baseline, color, label) in enumerate(((False, RED, "no baseline"), (True, BLUE, "running-average baseline"))):
1448        finals = [train_reinforce(Bandit(WIN_CHANCES, offset=5.0, seed=s), 400, 0.1, baseline=baseline, seed=s)["probs"][-1][2] for s in range(20)]
1449        a2.scatter(finals, row + rng.uniform(-0.15, 0.15, len(finals)), color=color, alpha=0.8)
1450    a2.set_yticks([0, 1], ["no baseline", "baseline"])
1451    a2.set(xlabel="final probability of the best arm, C", title="20 runs each, 400 pulls", xlim=(-0.05, 1.05), ylim=(-0.6, 1.6))
1452    fig.tight_layout()
1453    figs["baselines"] = fig
1454
1455    # --- 3. PPO: the clipped objective and ratio drift -----------------------
1456    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.4))
1457    ratios = np.linspace(0.4, 1.8, 300)
1458    for adv, color in ((1.0, BLUE), (-1.0, RED)):
1459        a1.plot(ratios, ratios * adv, color=color, ls="--", alpha=0.5)
1460        a1.plot(ratios, ppo_clipped_objective(ratios, adv), color=color, label=f"advantage {adv:+.0f}")
1461    a1.axvspan(0.8, 1.2, color=MUTED, alpha=0.2, label="band 1 ± 0.2")
1462    a1.set(xlabel="probability ratio  π_new / π_old", ylabel="objective for one sample", title="The clip: flat outside the band")
1463    a1.legend(frameon=False, loc="upper left")
1464    drift = ppo_drift_experiment()
1465    a2.plot(drift["unclipped"].max(axis=1), color=RED, label="no clip")
1466    a2.plot(drift["clipped"].max(axis=1), color=BLUE, label="clip ε = 0.2")
1467    a2.axhline(1.2, color=MUTED, ls="--")
1468    a2.set(xlabel="pass over the same 16 pulls", ylabel="largest ratio in the batch", title="Reusing one batch 50 times")
1469    a2.legend(frameon=False)
1470    fig.tight_layout()
1471    figs["ppo_clip"] = fig
1472
1473    # --- 4. GRPO on the addition prompts ------------------------------------
1474    run = train_grpo(steps=60, seed=0)
1475    fig, ax = plt.subplots(figsize=(6.5, 3.6))
1476    ax.bar(np.arange(1, 61), run["no_signal"], color=MUTED, alpha=0.6, label="prompts whose group all agreed (no signal)")
1477    ax.plot(np.arange(61), run["accuracy"], color=BLUE, lw=2, label="chance of a correct answer")
1478    ax.set(xlabel="GRPO step (8 answers per prompt)", ylabel="share", title="GRPO with a verifier: 8 addition prompts", ylim=(0, 1.35))
1479    ax.set_yticks(np.linspace(0, 1, 6))
1480    ax.legend(frameon=False, loc="upper left")
1481    figs["grpo"] = fig
1482
1483    # --- 5. The flawed reward model -----------------------------------------
1484    fig, ax = plt.subplots(figsize=(6.5, 3.6))
1485    fine = np.linspace(1, 10, 200)
1486    ax.axvspan(min(RATED_LENGTHS), max(RATED_LENGTHS), color=MUTED, alpha=0.2, label="lengths people rated")
1487    ax.plot(fine, true_quality(fine), color=BLUE, label="true quality")
1488    ax.plot(fine, proxy_reward(fine), color=RED, label="reward model (a fitted line)")
1489    ax.plot(RATED_LENGTHS, true_quality(np.array(RATED_LENGTHS)), "o", color=BLUE)
1490    ax.axhline(0, color="#4b5563", lw=0.8)
1491    ax.set(xlabel="answer length (sentences)", ylabel="score", title="The reward model extrapolates; the truth turns over")
1492    ax.legend(frameon=False, loc="lower left")
1493    figs["length_rewards"] = fig
1494
1495    # --- 6. Reward hacking during training, and the leash sweep --------------
1496    fig, (a1, a2) = plt.subplots(1, 2, figsize=(9, 3.6))
1497    for label, reward_fn, beta, color in (
1498        ("flawed reward, no leash", proxy_reward, 0.0, RED),
1499        ("flawed reward, KL leash β = 0.3", proxy_reward, 0.3, ORANGE),
1500        ("verifiable reward (the true goal)", true_quality, 0.0, BLUE),
1501    ):
1502        a1.plot(optimise_lengths(reward_fn, beta=beta)["true"], color=color, label=label)
1503    a1.set(xlabel="training step", ylabel="true quality of the policy", title="Rise, then fall: reward hacking", ylim=(0.6, 1.02))
1504    a1.legend(frameon=False, loc="lower left", fontsize=8)
1505    rows = leash_sweep(betas=np.geomspace(0.04, 20, 40))
1506    kls = [r["kl"] for r in rows]
1507    a2.plot(kls, [r["proxy"] for r in rows], color=RED, label="reward model's score")
1508    a2.plot(kls, [r["true"] for r in rows], color=BLUE, label="true quality")
1509    a2.set_xscale("log")
1510    a2.set(xlabel="KL from the reference (looser leash →)", ylabel="average score", title="Best policy for each leash strength")
1511    a2.legend(frameon=False, loc="lower left")
1512    fig.tight_layout()
1513    figs["reward_hacking"] = fig
1514
1515    return figs

Plot this lesson's data. matplotlib is imported here, and only here, so the lesson itself needs nothing beyond NumPy.

def demo() -> None: on GitHub
1523def demo() -> None:
1524    banner("1. A bandit: three slot machines, one hidden best")
1525    say(
1526        """
1527        Arms A, B and C pay 1 with chance 20%, 50% and 80%, else 0. The agent
1528        is not told this. A policy that picks uniformly earns 0.5 per pull on
1529        average; always picking C earns 0.8.
1530        """
1531    )
1532    table(
1533        ["policy", "expected reward J"],
1534        [("uniform", expected_reward(np.full(3, 1 / 3), np.array(WIN_CHANCES))), ("always C", expected_reward(np.array([0, 0, 1.0]), np.array(WIN_CHANCES)))],
1535        floatfmt=".2f",
1536    )
1537
1538    banner("2. REINFORCE: one step, by hand")
1539    say("Uniform logits (0, 0, 0). The agent pulls C and wins: reward 1. Step = 0.5 × 1 × (one-hot − π).")
1540    table(["arm", "∇ log π(C)", "new probability"], zip(ARMS, grad_log_prob(np.zeros(3), 2), worked_reinforce_step()), floatfmt=".3f")
1541    history = train_reinforce(Bandit(WIN_CHANCES, seed=0), steps=500, lr=0.1)
1542    say(
1543        f"""
1544        Now 500 pulls. Reward over the first 100: {history['rewards'][:100].mean():.2f};
1545        over the last 100: {history['rewards'][-100:].mean():.2f}. Final policy:
1546        """
1547    )
1548    table(["arm", "win chance", "final probability"], zip(ARMS, WIN_CHANCES, history["probs"][-1]), floatfmt=".3f")
1549    takeaway("Step along reward × ∇ log π(action): rewarded actions become more likely, and on average the best arm wins.")
1550
1551    banner("3. Baselines: same average gradient, far less noise")
1552    for b in (0.0, 11.0):
1553        stats = gradient_estimate_stats(np.zeros(2), np.array([10.0, 12.0]), baseline=b)
1554        say(f"Rewards 10 and 12, baseline {b:g}: mean gradient for arm 2 = {stats['mean'][1]:+.2f}, variance = {stats['variance'][1]:.2f}")
1555    wrong = {
1556        bl: np.mean([train_reinforce(Bandit(WIN_CHANCES, offset=5.0, seed=s), 400, 0.1, baseline=bl, seed=s)["probs"][-1][2] < 0.5 for s in range(20)])
1557        for bl in (False, True)
1558    }
1559    say(
1560        f"""
1561        Offset every reward by 5. Of 20 training runs, {wrong[False]:.0%} lock onto
1562        a worse arm without a baseline, and {wrong[True]:.0%} with a running-average baseline.
1563        """
1564    )
1565    takeaway("Subtract the typical reward: the advantage says 'better or worse than usual', which is the only thing that matters.")
1566
1567    banner("4. PPO: the probability ratio and the clip")
1568    table(
1569        ["ratio", "advantage", "ratio × A", "clipped objective"],
1570        [(r, a, r * a, ppo_clipped_objective(r, a)) for r, a in ((1.1, 3.0), (1.5, 2.0), (0.5, -1.0), (1.5, -1.0))],
1571        floatfmt=".2f",
1572    )
1573    drift = ppo_drift_experiment()
1574    say(
1575        f"""
1576        One batch of 16 pulls, reused for 50 passes. Without the clip the largest
1577        ratio reaches {drift['unclipped'].max():.2f} and C's probability {drift['unclipped_probs'][2]:.2f}
1578        from sixteen pulls. With it, the ratio stops at {drift['clipped'].max():.2f} and C
1579        sits at {drift['clipped_probs'][2]:.2f}.
1580        """
1581    )
1582    say(f"KL leash: reward 1.0 for an answer at probability 0.6 vs reference 0.3, β = 0.1 → {kl_penalised_reward(1.0, np.log(0.6), np.log(0.3), 0.1):.3f}")
1583    takeaway("PPO reuses expensive samples for several steps, and the clip keeps each batch from pulling the policy too far.")
1584
1585    banner("5. GRPO: advantages from a group of answers, no value network")
1586    table(
1587        ["group rewards", "advantages"],
1588        [(str(r), np.array2string(group_advantages(np.array(r, dtype=float)), precision=2)) for r in ([1, 0, 0, 1], [1, 0, 0, 0], [1, 1, 1, 1])],
1589    )
1590    run = train_grpo(steps=60, seed=0)
1591    say(
1592        f"""
1593        GRPO with a verifier on 8 addition prompts, 8 answers each: accuracy goes from
1594        {run['accuracy'][0]:.0%} to {run['accuracy'][-1]:.0%} in 60 steps. By the end,
1595        {run['no_signal'][-10:].mean():.0%} of groups are unanimous and teach nothing.
1596        """
1597    )
1598    takeaway("Compare each answer with its siblings: the group mean is the baseline, and a verifier is the reward.")
1599
1600    banner("6. Reward hacking: a reward model that loves length")
1601    intercept, slope = fit_reward_model()
1602    say(f"Fitted on answers of 1 to 4 sentences, the reward model is {intercept:.4f} + {slope:.4f} × sentences.")
1603    table(
1604        ["sentences", "true quality", "reward model"],
1605        [(int(n), true_quality(n), proxy_reward(n)) for n in (1, 2, 3, 4, 6, 8, 10)],
1606        floatfmt=".2f",
1607    )
1608    free, leashed, true_run = optimise_lengths(), optimise_lengths(beta=0.3), optimise_lengths(true_quality)
1609    table(
1610        ["training against", "true quality: start", "peak", "end"],
1611        [
1612            (name, r["true"][0], r["true"].max(), r["true"][-1])
1613            for name, r in (("flawed reward", free), ("flawed reward + KL β=0.3", leashed), ("verifiable reward", true_run))
1614        ],
1615        floatfmt=".2f",
1616    )
1617    say(
1618        f"""
1619        Against the flawed reward, quality rises while lengthening helps, then falls below
1620        where it started as the policy settles on {SENTENCES[free['probs'].argmax()]}-sentence answers,
1621        even as the reward model's score climbs from {free['proxy'][0]:.2f} to {free['proxy'][-1]:.2f}.
1622        """
1623    )
1624    takeaway(
1625        "The policy optimises the reward you wrote, not the goal you meant. Prefer verifiable rewards, "
1626        "keep a KL leash, and watch a held-out measure of the real goal."
1627    )