An annotated companion · AI Primer

Scaling Laws for Reward Model Overoptimization, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is distributed under arXiv's standard non-exclusive licence, which does not grant permission to republish its figures or tables, so its figures are redrawn from scratch and only a few selected numbers appear in tables written for this page, each attributed. The paper's figures are raster images; the coefficients used to redraw its curves were read off the plotted points of its Figure 3 and cross-checked against its Figures 1 and 12, so they are good to about the second decimal place and are labelled as read off the plot. The equations are reproduced with every symbol decoded (mathematics is not copyrightable); the peak formulas and a few worked illustrations are this page's own and say so. Numbers marked illustrative are made up for teaching. Read the original alongside: every section links to it.

How to read this page

  • Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
  • The redrawn Figure 1 lets you pick a reward model size and watch where the true score peaks. Try it: the two laws lets you bend the two curves the paper fits, and the noise explorer shows why one kind of Goodhart can never make the true score fall.

Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The paper builds on RLHF as the InstructGPT companion describes it, and on PPO for the reinforcement learning. The alignment lesson draws Goodhart's law as two curves, and the reinforcement learning lesson builds a flawed reward model from scratch and watches a policy exploit it.

Abstract

“Because the reward model is an imperfect proxy, optimizing its value too much can hinder ground truth performance, in accordance with Goodhart's law.”Gao, Schulman & Hilton (2022), Abstract. Read the original

Everyday picture

A student preparing for an exam practises on old papers marked by a teaching assistant who is good but not perfect. Early on, pleasing the assistant and learning the subject are the same thing. Push hard enough and the student starts learning the assistant's quirks instead: the phrasings it rewards, the length it likes. Their marks keep rising while their real understanding stalls, then slides. Everyone has seen this happen; nobody had measured how fast it happens, or how a better assistant changes it.

What the paper claims

  • A clean, cheap experiment. A large “gold” reward model stands in for people, so the true score can be computed for every output.
  • The true score rises, then falls, in a predictable shape as a policy is optimized against a smaller proxy reward model. The shape differs between best-of-n sampling and reinforcement learning, and each has a two-number formula that fits it.
  • The two numbers change smoothly with reward model size, so a small experiment predicts a bigger one.
  • Bigger policies start higher but overoptimize about as much; a KL penalty in their setup acts like early stopping rather than a cure.

Why it matters today

Every system tuned against a learned judge, whether a reward model, a preference model or another model grading answers, faces this curve. The paper turned “don't optimize too hard” from folklore into a measurable quantity, and its “gold model as stand-in for people” trick was borrowed by later work, such as Let's Verify Step by Step, to study reward models cheaply.

1 Introduction · original

“Optimizing too much against such a model eventually hinders the true objective, a phenomenon we refer to as overoptimization.”Gao, Schulman & Hilton (2022), §1

Everyday picture

Goodhart's law: once a measure becomes a target, it stops being a good measure. In RLHF the measure is a reward model, a network trained to predict which of two answers a person would prefer, and the target-chaser is the language model being tuned. The paper calls the damage overoptimization.

Tiny example: how far has the policy moved?

To draw “true score against how hard we optimized”, the paper needs a ruler for “how hard”. It uses how far the tuned policy has drifted from where it started, measured by KL divergence, in nats. Then it takes the square root, because KL grows like a squared distance. A policy 4 nats away is at distance d = 2; one 9 nats away, d = 3.

In words: “the distance d is the square root of the KL divergence from the optimized policy to the initial policy.”

With the numbers: a KL of 4 nats gives d = 2; the largest best-of-n run in the paper, about 10 nats, gives d = 3.16; a reinforcement learning run at 100 nats gives d = 10.

In Python:

import math
# d = √KL for three policies, KL in nats
[round(math.sqrt(KL), 2) for KL in (4, 10, 100)]  # → [2.0, 3.16, 10.0]

The two laws

The main result is two formulas for the gold (true) score R as a function of d, one for each way of optimizing. Both start at zero, because every score is measured relative to the starting policy.

In words: “for best-of-n, the gold score is the distance times (a gain per unit distance minus a penalty that grows with distance).”

With the numbers: for the smallest proxy (3M parameters), the coefficients read off the paper's Figure 3 are α ≈ 0.49 and β ≈ 0.12. At n = 1,000 samples the KL is 5.91 nats (§2 shows why), so d = 2.43 and R = 2.43 × (0.49 − 0.12 × 2.43) = 0.48. Figure 1(a) shows the 3M gold curve at about the same height there.

In Python:

import math
# α_bon, β_bon for the 3M proxy, read off Figure 3
alpha, beta = 0.49, 0.12
# d for best-of-1000 (KL = log n − (n − 1)/n, see §2)
d = math.sqrt(math.log(1000) - 999 / 1000)
round(d, 2)  # → 2.43
# R_bon(d) = d (α − β d)
round(d * (alpha - beta * d), 3)  # → 0.482

In words: “for reinforcement learning, the gold score is the distance times (a gain per unit distance minus a penalty that grows with the logarithm of the distance).”

With the numbers (illustrative): the paper plots but does not print its RL coefficients, so take α = 0.35 and β = 0.15, chosen to resemble the 3M curve in Figure 1(b). At KL = 25 nats, d = 5 and R = 5 × (0.35 − 0.15 × ln 5) = 5 × (0.35 − 0.241) = 0.543.

In Python:

import math
# α_RL, β_RL (illustrative)
alpha, beta = 0.35, 0.15
d = math.sqrt(25)
# R_RL(d) = d (α − β log d), natural log
round(d * (alpha - beta * math.log(d)), 3)  # → 0.543

The paper notes the RL form has an infinite slope at d = 0, so it probably fails very close to the start; it tried smoother forms and found they fitted and extrapolated worse (Appendix B).

Reading it: this is Figure 1(a) of the paper rebuilt from its fitted law. The x-axis is KL distance from the starting policy, drawn on a square-root scale as in the paper, so equal steps are equal steps in d; the y-axis is the gold score. Each line is the best-of-n gold curve R = d(α − βd) for one proxy size, using α and β read off the paper's Figure 3: the smallest (3M), the largest (3B), and the size the slider picks. Every curve climbs at first, when pleasing the proxy and pleasing the gold model are the same thing, then bends over and falls as the proxy's mistakes take over. Slide the size up and the peak moves up and to the right: a bigger proxy can be pushed further before it misleads. The paper measured up to KL ≈ 10 nats; the lines continue to 16 as the law's forecast, which is how the 3B peak (about 13 nats) shows. The readout gives the chosen size's coefficients and its peak.

Four findings, in one line each

  • RL against best-of-n: measured in KL, RL is slower at both optimizing and overoptimizing, which says KL is a poor common yardstick (§3.5, §4.1).
  • Smooth scaling: α and β move smoothly with proxy size, roughly linearly in its logarithm (§3.2).
  • Policy size matters little to how much overoptimization happens (§3.4).
  • A KL penalty did not move the gold-score frontier in their setup (§3.6).

Why it matters

A scaling law turns a qualitative warning into a forecast: from small, cheap runs you can predict how far a larger run can be pushed before the true score turns down. The goodhart_curve in the alignment lesson has exactly this rise-then-fall shape, built from a proxy that counts padding as content.

Try it: the two laws · original

Everyday picture

Two hikers climb the same hill in different ways. One takes long careful strides and turns back sharply after the summit; the other starts steeply and then drifts down a long, gentle slope. Same idea (up, then down), different shapes.

Tiny example

With α = 0.35 and β = 0.15 (illustrative) for both laws, the best-of-n curve peaks at d = α / 2β = 1.17 (KL 1.4 nats) and crosses zero at d = α / β = 2.33. The RL curve peaks much later, at d = eα/β − 1 = 3.79 (KL 14.4), and crosses zero at d = eα/β = 10.3 (KL 106).

In words: “the best-of-n law peaks where the distance equals the gain divided by twice the penalty; the RL law peaks where the log of the distance equals gain over penalty, minus one.” (These follow by setting each law's slope to zero; the paper uses the first in its Figure 12.)

With the numbers (illustrative): α = 0.35, β = 0.15 give d* = 0.35 / 0.30 = 1.17 for best-of-n and d* = e1.333 = 3.79 for RL. The peak heights are α² / 4β = 0.204 and β d* = 0.569.

In Python:

import math
alpha, beta = 0.35, 0.15
# best-of-n: d* = α / 2β, and the peak score α² / 4β
d_bon = alpha / (2 * beta)
round(d_bon, 2), round(d_bon ** 2, 1), round(alpha ** 2 / (4 * beta), 3)  # → (1.17, 1.4, 0.204)
# RL: d* = e^(α/β − 1), and the peak score β d*
d_rl = math.exp(alpha / beta - 1)
round(d_rl, 2), round(d_rl ** 2, 1), round(beta * d_rl, 3)  # → (3.79, 14.4, 0.569)

Try it: drag α and β. Raising α lifts and delays both peaks; raising β pulls both down sooner. Notice that the RL curve always falls more slowly after its peak: the logarithm grows far more slowly than d itself.

Reading it: both curves use the same α and β, which the paper never does (each method gets its own fit); sharing them isolates the difference in shape. The x-axis is d, the square root of KL; the y-axis is the gold score. The solid line is the best-of-n law, a parabola in d that turns over quickly and falls steeply. The dashed line is the RL law, which climbs more steeply at first, peaks later and declines slowly. The readout gives each peak. Every number here is illustrative.

Why it matters

Knowing the shape lets you find the peak from a handful of early measurements, which is the whole use of a scaling law: stop before the turn, not after it.

2 Methodology · original

“In BoN, we generate n trajectories for the policy and use the reward model to pick the one with the highest proxy RM score.”Gao, Schulman & Hilton (2022), §2

Everyday picture

Two ways to get better essays from the same student. Best-of-n: have them write n drafts and hand in whichever the marker likes most; the student never changes. Reinforcement learning: after each batch of essays, nudge the student's habits towards what the marker liked; the student changes a little every round.

The setting

  • The InstructGPT environment: prompts are natural-language instructions, the policy writes a response, a reward model scores it. See the InstructGPT companion.
  • Every policy starts from a pretrained GPT-3-series model given two epochs of supervised fine-tuning on InstructGPT's human demonstrations. Every reward model is the same architecture with a single-number output head.
  • RL uses PPO with the KL penalty set to zero everywhere except §3.6, and mostly default hyperparameters.
  • Best-of-n scores are computed with an unbiased estimator (from WebGPT) rather than by literally resampling, which gives every n from 1 up to the maximum at once.

Tiny example: how far does best-of-n move the policy?

Picking the best of 2 is a gentle nudge; picking the best of 1,000 is a hard one. Treating “best of n” as a new policy, its KL from the original has an exact formula (from Stiennon et al., 2020). For n = 2 it is ln 2 − 1/2 = 0.19 nats.

In words: “the KL divergence of best-of-n from the original policy is the natural log of n, minus (n minus one) over n.”

With the numbers: n = 2 gives 0.19 nats; n = 1,000 gives 6.91 − 1.00 = 5.91; n = 60,000 gives 11.00 − 1.00 = 10.00. Going from 1,000 to 60,000 samples, sixty times the compute, buys only 4 more nats: best-of-n moves the policy very slowly.

In Python:

import math
def KL_bon(n):
    # log n − (n − 1)/n, in nats
    return math.log(n) - (n - 1) / n
[round(KL_bon(n), 2) for n in (2, 1000, 60000)]  # → [0.19, 5.91, 10.0]

Why it matters

Best-of-n is the simplest possible optimizer and needs no training, which is why it is also used at inference time (the Training Verifiers companion uses it to pick maths solutions). Its KL has a closed form; RL's does not and must be measured.

2.1 Synthetic data setup · original

“We instead use a synthetic task where the ground truth is defined to be the output of a particular large ‘gold’ RM.”Gao, Schulman & Hilton (2022), §2.1

Everyday picture

To study how well trainee judges copy a head judge, you don't need the public at all: let the head judge label thousands of pairs, train trainees on those labels, and then ask how badly things go when contestants play to a trainee. Because the head judge is always available, you can check every contestant against the “truth” for free.

Real RLHF Labellers Humancomparisons Proxy RM(learned) This paper's synthetic setup Policy outputpairs Gold RM(6B, fixed) Syntheticcomparisons Proxy RM3M to 3B RL or best-of-n outputs go back to the gold RM to be scored

Hover or tap a block. Start with Gold RM: it replaces the labellers.

The paper's Figure 2, redrawn, with the optimization step and the gold model's second job (scoring what the optimization produces, dashed) added from the text of §2 and §3. Based on Gao, Schulman & Hilton (2022), Figure 2.

Reading it: the top row is ordinary RLHF: people compare pairs of answers, and those comparisons train a proxy reward model. The bottom row swaps the people for a fixed 6-billion-parameter reward model (the gold model, taken from InstructGPT). It compares pairs of policy outputs and writes the synthetic comparisons that train each proxy. The proxy is then optimized against, by RL or best-of-n. The dashed wire, from the optimization back up to the gold model, is the gold model's second job: it scores whatever the optimization produces, giving a true score for free. In both rows the proxy only ever sees labels, never the thing producing them, which is what makes the synthetic setup a fair stand-in for the real one.

The numbers

  • Proxy reward models from 3M to 3B parameters. Two even smaller ones were trained, landed near chance accuracy, and were dropped.
  • 100,000 synthetic comparisons, 10% held out to measure validation loss.
  • Labels are hard: the answer with the higher gold score always wins. Sampling labels from the gold model's probabilities gave noisier results.

Why it matters

Human preference data is expensive and noisy, and a scaling law needs many runs with small effects. A gold model makes the experiment repeatable. The price, which §4.5 names, is that it can only study the gap between the proxy and its labels, not the gap between the labels and what people really want. Let's Verify Step by Step later borrowed the same move, using a large reward model to label data for small ones.

2.2 Recalibration · original

Everyday picture

Two thermometers, one in Celsius and one in Fahrenheit, can both be right and still disagree about the number. To compare reward models, first put them on the same scale.

Tiny example

A reward model scores the starting policy's answers 2.1, 1.9 and 2.3 on average 2.1 (illustrative). Subtract 2.1 from every score, and the starting policy sits at 0. A second model that averaged −0.4 gets −0.4 subtracted. Now “+0.5” means the same thing for both: half a unit better than where we started.

In Python:

# scores of the starting policy's answers (illustrative)
scores = [2.1, 1.9, 2.3]
mean = sum(scores) / len(scores)
# recentre so the starting policy averages 0
[round(s - mean, 2) for s in scores]  # → [0.0, -0.2, 0.2]

What the paper does

  • A reward model's scores only matter up to a constant (adding 5 to every score changes no comparison), so each is shifted to make the initial policy average 0. That is why every curve on this page starts at 0.
  • The gold scores are also scaled to unit variance (a footnote says this turned out unnecessary).
  • Because the hard labels ignore the gold model's confidence, each proxy's scores are rescaled to best predict a set of soft labels. All of this is applied after the runs: it cannot change which sample best-of-n picks, and probably doesn't change RL, since the Adam optimizer ignores the scale of the loss.

Why it matters

Without a shared zero, “the gold score went up by 0.5” would mean different things in different runs, and no curve could be compared with another.

3 Results · original

Everyday picture

Once the experiment is cheap, you can turn one dial at a time: how big the proxy is, how much data it saw, how big the policy is, which optimizer is used, and whether there's a leash. Each subsection below turns one dial.

Tiny example

Holding everything else fixed (policy 1.2B, 90,000 training comparisons), only the proxy size changes in §3.2; only the data size in §3.3 (proxy 12M); only the policy size in §3.4.

Why it matters

Changing one thing at a time is what lets the paper say which dial matters and which doesn't.

3.1 Fitting and validating the functional forms · original

“As this experiment was conducted after the functional form was hypothesized based on data up to 6 nats, this was a true advance prediction.”Gao, Schulman & Hilton (2022), §3.1

Everyday picture

Anyone can draw a curve through points they have already seen. The test of a law is whether it predicts points nobody has measured yet.

Tiny example: the advance prediction

The best-of-n form was chosen using runs up to n = 1,000, which is 5.91 nats. Then the team ran n = 60,000 (10.0 nats) and checked the fitted curves against the new data. The KL formula gives both numbers.

One question, many samples

Averaged over many prompts the curves are smooth; on a single prompt they are not. The paper's Table 2 follows one riddle, “What is full of holes but still holds water?”, through best-of-n with a 1.2B policy and a 12M proxy:

Selected rows of Table 2 of Gao, Schulman & Hilton (2022), answers shortened.
nChosen answerProxyGold
1a rambling answer about mussels and clams−0.19−0.52
10“A sponge”0.230.48
100a description of a tornado0.90−0.34
30,000a description of a pothole0.950.55

Reading it: read down the rows as optimization gets harder. The proxy score climbs every time, because best-of-n always picks the proxy's favourite. The gold score does not: “A sponge” at n = 10 is the riddle's answer and scores well; at n = 100 the proxy prefers a confident paragraph about tornadoes, which the gold model rates below the starting point. That row is overoptimization in miniature: a higher proxy score bought with a worse answer.

What could not be fitted

The team also tried to model the proxy score itself and could not get a satisfying fit; for best-of-n, a straight line in d looked right but fitted poorly. The gold scores follow the laws; the proxy scores do not follow any simple one found here.

Why it matters

A law chosen before the test data existed, and then confirmed on it, is much more believable than one fitted afterwards. The same discipline (fix the rule, then measure) is behind the release gates in the alignment lesson's release_gate.

3.2 Scaling with reward model size · original

Everyday picture

A more experienced marker is harder to fool: you have to drift further from honest work before their marks stop tracking quality, and the best marks you can honestly earn from them are higher.

Tiny example

From the coefficients read off Figure 3, the 3M proxy's best-of-n curve peaks at KL ≈ 4.2 with gold score 0.50. The 3B proxy's peaks at KL ≈ 13.4 with gold score 1.21: more than twice as high, three times as far out.

Coefficients of the gold-score laws for each proxy size (policy 1.2B, 90,000 comparisons). Read off the plotted points of Figure 3 of Gao, Schulman & Hilton (2022), to about ±0.005; the peak columns are computed on this page.
Proxy sizeαbonβbonβRLpeak KL (bon)peak gold (bon)

Reading it: each row is one proxy. Down the table, αbon (how fast the gold score rises at first) grows from 0.49 to 0.66 while βbon (how fast the proxy's mistakes pull it back down) shrinks from 0.120 to 0.090. For RL the paper found it could hold αRL fixed across sizes, leaving βRL to carry the whole trend, falling from 0.175 to about 0.12. The last two columns follow from the first two: the peak moves out and up with every step in size.

In words: “the best-of-n gold curve peaks at distance α over 2β, and its highest value is α squared over 4β.”

With the numbers: 3M: d* = 0.49 / 0.24 = 2.04, so KL* = 4.2 and R* = 0.2401 / 0.48 = 0.50. 3B: d* = 0.66 / 0.18 = 3.67, KL* = 13.4, R* = 0.4356 / 0.36 = 1.21.

In Python:

# (α_bon, β_bon) read off Figure 3, for the 3M and 3B proxies
coefs = [(0.49, 0.12), (0.66, 0.09)]
# for each: d* = α / 2β, KL* = d*², R* = α² / 4β
[(round(a / (2 * b), 2), round((a / (2 * b)) ** 2, 1), round(a ** 2 / (4 * b), 2)) for a, b in coefs]  # → [(2.04, 4.2, 0.5), (3.67, 13.4, 1.21)]

The trend in size

The paper describes the coefficients as following “approximate logarithmic trends”: a straight line against the logarithm of the parameter count. Fitting a line through the nine read-off values gives αbon ≈ 0.086 + 0.063 log10N and βbon ≈ 0.192 − 0.011 log10N (this page's fit).

In Python:

import math
N = [3e6, 12e6, 25e6, 42e6, 85e6, 300e6, 680e6, 1.2e9, 3e9]
alpha = [0.49, 0.51, 0.56, 0.55, 0.60, 0.63, 0.65, 0.65, 0.66]
x = [math.log10(v) for v in N]
x_bar, a_bar = sum(x) / len(x), sum(alpha) / len(alpha)
# least-squares slope and intercept of α against log10 N
slope = sum((xi - x_bar) * (ai - a_bar) for xi, ai in zip(x, alpha)) / sum((xi - x_bar) ** 2 for xi in x)
round(a_bar - slope * x_bar, 3), round(slope, 3)  # → (0.086, 0.063)

A slip in the appendix

The paper's Figure 12 is titled “Max BoN gold scores (αbon/2βbon)” and plots values from about 1.4 to 3.6. But α/2β is the distance d* at which the gold score peaks, not the peak score (that is α²/4β, which never exceeds about 1.2 here, as Figure 1 shows). The plotted values match d* for every size to within about 0.05, so the figure is right about what it computes and mislabelled about what it shows. It also includes proxies smaller than 3M, the ones §2.1 says were excluded.

Why it matters

Since the peak grows smoothly with the proxy's size, you can forecast how hard a given reward model can safely be pushed without running the expensive experiment. For the proxy scores the fits are less trustworthy, and extrapolated to larger KL they come out systematically too low: the proxy score eventually grows roughly in a straight line in d.

3.3 Scaling with reward model data · original

Everyday picture

A trainee judge who has seen twenty contests has learned almost nothing; after a few thousand they start to get it. Showing them the same twenty contests four times over does not help.

Tiny example: near-chance loss

A reward model that knows nothing gives every pair 50/50, and its cross-entropy loss is then ln 2 = 0.693 per comparison. The paper's Figure 13 reports three settings, averaged across reward model sizes: 2,000 comparisons once (loss 0.6861), 2,000 comparisons four times (0.6839), and 8,000 comparisons once (0.6549).

In words: “a reward model that always says fifty-fifty pays minus the log of one half, which is the log of 2, on every comparison.”

With the numbers: ln 2 = 0.693. The 2,000-comparison run is only 0.007 below it, four epochs on the same data only 0.009, but 8,000 fresh comparisons 0.038: four times the unique data bought four times the improvement, while four times the passes bought almost nothing.

In Python:

import math
L_chance = math.log(2)
round(L_chance, 3)  # → 0.693
# validation losses from Figure 13: 1 × 2000, 4 × 2000, 1 × 8000 comparisons
losses = {"1x2000": 0.686109, "4x2000": 0.683869, "1x8000": 0.654857}
{k: round(L_chance - v, 3) for k, v in losses.items()}  # → {'1x2000': 0.007, '4x2000': 0.009, '1x8000': 0.038}

What the paper finds

  • Below about 2,000 comparisons, reward models of every size barely beat chance, and the gold scores after optimization show it.
  • Past that threshold, more data means better gold scores and less goodharting, and larger reward models improve faster, though the threshold itself does not move earlier for them (a footnote says this contradicts other internal findings).
  • Weak evidence for a tidy idea: two reward models with the same validation loss are about equally robust, however that loss was reached.

Why it matters

Unique preference data is what buys robustness, and there is a minimum below which a reward model is not worth optimizing against at all. The reward_model_loss in the training stages lesson is the loss being measured here.

3.4 Scaling with policy size · original

Everyday picture

A stronger student starts closer to full marks, so has less to gain from gaming the marker. You might expect them to find the marker's blind spots faster too. In these experiments they don't.

Tiny example

Two policies, 1.2B and 6B, each optimized against the same 12M proxy (and again against a 3B proxy). The 6B policy's gold score rises less from its start to its peak, yet the two peak at almost the same KL, and the gap between proxy and gold score, a measure of how much the proxy is being exploited, is almost the same for both.

Why it matters

If exploitation doesn't speed up with policy size, then the reward model, not the policy, sets how far you can push. Only two policy sizes were tried, which the paper lists among its limitations; §4.4 offers a hypothesis.

3.5 RL versus best-of-n · original

Everyday picture

Best-of-n only picks among what the original student would already write, so it stays close to home. RL changes the student a little every step, and those little changes add up to a long way from home.

Tiny example: KL spent

Best-of-60,000 reaches only 10 nats. The RL runs in Figure 1(b) travel to about 100 nats: ten times further. The paper explains why: best-of-n's KL grows like log n (the formula), while an RL policy keeps moving from wherever the last step left it, and without a penalty its KL grows roughly with the square of the number of steps.

Proxy score as the ruler

If you put the proxy score on the x-axis instead of KL (the paper's Figure 8), best-of-n and RL look much more alike: the gold score first rises with the proxy score, then peels away from it. Two differences remain: RL starts with a bigger proxy-to-gold gap, and yet peaks at a higher gold score than best-of-n does.

Why it matters

KL is fine for comparing runs of one method and misleading across methods; §4.1 spells out why.

3.6 Effect of the KL penalty · original

“The KL penalty only causes the gold RM score to converge earlier, but does not affect the KLRL-gold reward frontier, and so the effect of the penalty on the gold score is akin to early stopping.”Gao, Schulman & Hilton (2022), §3.6

Everyday picture

A dog on a leash can't wander far. But if the leash only stops the dog where it would have been anyway at that length of walk, then the leash is just a shorter walk.

Tiny example

Take five RL runs with penalty coefficients 0, 0.01, 0.05, 0.1 and 0.5 (their Figure 9, 1.2B policy and 1.2B proxy). Plotted against KL, the five gold curves lie on top of each other; a bigger penalty just makes a run stop sooner along the shared curve. The proxy scores, though, are higher at a given KL with a penalty: the penalty buys more proxy reward per nat, but not more gold reward, so the proxy-gold gap is strictly larger. That is why every other RL run in the paper uses no penalty.

A puzzle left open

PPO's clipped objective already restrains each update relative to the previous policy, and this slows the growth of KL from the initial policy. The paper does not know why that indirect effect appears to cause less overoptimization than an explicit penalty, and warns the whole result may be sensitive to hyperparameters.

Why it matters

The penalty is standard in RLHF (the InstructGPT companion decodes it, and the kl_penalised_reward in the reinforcement learning lesson builds it). This result says that, at least here, it works by stopping the run early, not by steering the policy somewhere safer. The reinforcement learning lesson's leash_sweep shows the same trade in a toy: each leash strength lands at a different point on one proxy-versus-truth curve.

4 Discussion · original

Everyday picture

Measurements first, meaning second. The discussion asks what the curves say about the ruler (KL), about why proxies fail, about repeated rounds of RLHF, and about what this experiment cannot see.

Tiny example

The two coefficients get an interpretation: α measures one kind of Goodhart (noise), β another (the proxy failing off its training data). §4.2 builds both.

Why it matters

A curve that fits is useful; a curve whose terms mean something tells you what to fix.

4.1 KL as a measure of optimization · original

Everyday picture

Miles on a car's odometer measure how far it went, not how much closer it got to where you wanted to go. Driving in circles adds miles too.

Tiny example

A policy could change the punctuation style of every answer: a real KL cost with no effect on either reward. Or it could change one crucial word in a few answers: a tiny KL with a big effect on behaviour. KL counts both the same way.

Why it matters

Within one method KL gives clean trends and consistent peak locations; across methods it does not measure the same thing, so “RL overoptimizes at higher KL than best-of-n” is a statement about the ruler as much as about the methods.

4.2 Four kinds of Goodhart · original

“Intuitively, we can interpret eq. 1 as stating that the optimization power expended is divided between optimizing the gold reward and selecting on the noise proportional to their variances.”Gao, Schulman & Hilton (2022), §4.2.1

Everyday picture

The paper borrows a four-way taxonomy (Manheim and Garrabrant, 2018) of how a measure can come apart from its goal.

Regressionalproxy = truth + noiseα term Extremaloff the training dataβ term Causalcorrelated, not causinglooks like regressional Adversarialpolicy games the proxynot seen here

Hover or tap a box to see how the paper maps it onto its curves.

The four Goodhart categories as §4.2 discusses them, drawn for this page; the paper has no figure for them. Based on Gao, Schulman & Hilton (2022), §4.2.

Reading it: the top row holds the two effects the paper ties to its coefficients. Regressional Goodhart (left) is the proxy being truth plus noise; the paper links it to α. Extremal Goodhart (right) is the proxy failing on inputs unlike its training data; the paper expects it to cause most of the downturn, the β term. Bottom left, causal Goodhart would look like regressional in these experiments. Bottom right, adversarial Goodhart, a policy deliberately manipulating the proxy, is drawn dashed because the paper's models are not capable enough to show it.

Tiny example: regressional Goodhart

Suppose the gold score X and the proxy's noise Z are both bell-curve distributed with average 0, X with spread (variance) 1 and Z with variance 0.5. You see an answer whose proxy score is 1.5. What gold score should you expect? Not 1.5: part of that 1.5 is probably luck in the noise. The formula below says 1.0.

In words: “the gold score you should expect, given a proxy score, is the average gold score plus the proxy's surprise (how far it is above what you'd expect) shrunk by the share of the proxy's spread that comes from real signal.”

With the numbers (illustrative): E[X] = E[Z] = 0, Var(X) = 1, Var(Z) = 0.5, x̂ = 1.5: 0 + (1.5 − 0 − 0) × 1 / 1.5 = 1.0. A third of the apparent gain was noise. A simulation of 200,000 random pairs, keeping those with proxy scores near 1.5, agrees.

In Python:

import random
EX, EZ, VarX, VarZ, x_hat = 0.0, 0.0, 1.0, 0.5, 1.5
# E[X | X̂ = x̂] = E[X] + (x̂ − E[X] − E[Z]) · Var(X) / (Var(X) + Var(Z)), with ε = 0 for Gaussians
EX + (x_hat - EX - EZ) * VarX / (VarX + VarZ)  # → 1.0
# check by simulation: draw gold X and noise Z, keep proxies X + Z close to 1.5
rng = random.Random(0)
kept = []
for _ in range(200000):
    X, Z = rng.gauss(0, 1), rng.gauss(0, 0.5 ** 0.5)
    if abs(X + Z - x_hat) < 0.05:
        kept.append(X)
round(sum(kept) / len(kept), 1)  # → 1.0

The paper proves this (Appendix A) when the noise is Gaussian (then ε = 0), and approximately when the noise is small and bounded (ε shrinks faster than Var(Z) as the noise's range goes to zero). The consequence: if noise were the only problem, the gold score would rise with the proxy score forever, just more slowly. The curves in Figure 1 turn down, so something else must be at work.

Extremal, causal, adversarial

  • Extremal. Optimization pushes answers away from anything the proxy was trained on, where its judgement weakens. If long answers were always better in its training data, it learns “longer is better”, which fails far outside it; the paper notes that optimized policies writing overly long answers is a real problem they have seen. The team expects this to cause most of the downturn, and reads the smooth fall in β with proxy size as proxies getting smoothly more robust. The fit_reward_model in the reinforcement learning lesson builds exactly this failure: a straight line fitted to short answers, extrapolated to long ones.
  • Causal. A feature correlated with quality because both have a common cause (length and informativeness, say). Selecting on the feature doesn't buy the quality. In these experiments it would look like regressional Goodhart.
  • Adversarial. A policy actively manipulating the proxy. The paper does not expect to see it with these models, and warns that if more capable systems do it, the scaling laws may break down.

Why it matters

The split says what each fix can reach. More data and averaging reduce noise (α); making reward models robust off their training distribution attacks the downturn (β). The alignment lesson's goodhart_curve is an extremal example: its proxy counts padding it was never taught to penalise.

Try it: pure noise never turns down

Everyday picture

Pick the tallest of n people using a tape measure that sometimes slips. The more people you measure, the taller your pick, on average, even though some of the “extra height” is slippage. A noisy ruler slows progress; it never reverses it.

Tiny example

Best of 16 with Var(X) = Var(Z) = 1: the best of 16 standard bell-curve draws averages 1.766 standard deviations above the mean. The proxy's spread is √2, so the chosen answer's proxy score averages 1.766 × √2 = 2.50, and its gold score 1.766 × 1 / √2 = 1.25: half the proxy's apparent gain, but a gain.

In words: “under pure noise, the gold score of the best-of-n pick is the expected best of n standard draws, times the gold variance, divided by the proxy's spread.” (This page's consequence of equation 1, not a formula from the paper.)

With the numbers: e16 = 1.766, Var(X) = Var(Z) = 1: 1.766 × 1 / 1.414 = 1.25, against a proxy score of 2.50.

In Python:

import math
def Phi(x):
    return 0.5 * (1 + math.erf(x / math.sqrt(2)))
def phi(x):
    return math.exp(-x * x / 2) / math.sqrt(2 * math.pi)
def e(n, steps=4000, lo=-8.0, hi=8.0):
    # expected maximum of n standard normal draws: ∫ x · n φ(x) Φ(x)^(n−1) dx
    h = (hi - lo) / steps
    return sum((lo + i * h) * n * phi(lo + i * h) * Phi(lo + i * h) ** (n - 1) * h for i in range(steps + 1))
e_16 = e(16)
round(e_16, 3)  # → 1.766
VarX, VarZ = 1.0, 1.0
# gold score of the pick, then its proxy score
round(e_16 * VarX / math.sqrt(VarX + VarZ), 2), round(e_16 * math.sqrt(VarX + VarZ), 2)  # → (1.25, 2.5)

Try it: drag the noise. With no noise the proxy and gold curves coincide; as noise grows the gold curve flattens, but it never bends back down, however far right you look.

Reading it: the x-axis is KL from the starting policy on a square-root scale, running from best-of-1 to best-of-100,000 through the KL formula of §2; the y-axis is score in units of the gold model's spread. The solid line is what the proxy reports for its favourite; the dashed line is that pick's gold score. Both rise without limit, and against d both are nearly straight lines: the best of n draws grows like √(2 log n), and d like √(log n). (The paper, too, finds its proxy scores grow roughly linearly in √KL.) The gap between them is regressional Goodhart; the paper reads α, the initial slope of the gold law, as the part of the optimization that survives it. Figure 1's gold curves turn down, which pure noise cannot do: that is the paper's evidence for the extremal kind.

Why it matters

If your reward model's errors were only random noise, optimizing harder would always help a little. The practical worry is the other kind, the errors that grow as the policy moves somewhere new.

4.3 Implications for iterated RLHF · original

Everyday picture

Instead of one long walk with one old map, walk a stretch, draw a fresh map of where you are, walk another stretch. Each map is only trusted near where it was drawn.

Tiny example

Online RLHF periodically collects fresh preferences on the current policy's outputs and trains a new reward model. Take the illustrative RL law (α = 0.35, β = 0.15) and a total distance d = 5 (KL 25). In one go, R = 0.543. Split the same distance into k = 4 legs, each with a freshly trained reward model: the formula below gives 1.583.

In words: “after k rounds, each covering a k-th of the distance with a fresh reward model, the gold score is the one-round law plus an extra β times d times the log of k.”

With the numbers (illustrative): k = 1 gives 5 × (0.35 − 0.15 × 1.609) = 0.543. k = 4 adds 0.15 × 5 × ln 4 = 1.040, for 1.583. The α term is untouched: fresh reward models do nothing about noise.

In Python:

import math
alpha, beta, d = 0.35, 0.15, 5.0
def R_iter(k):
    # d (α − β log d + β log k)
    return d * (alpha - beta * math.log(d) + beta * math.log(k))
round(R_iter(1), 3), round(R_iter(4), 3)  # → (0.543, 1.583)
# the gain from iterating: β d log k
round(beta * d * math.log(4), 3)  # → 1.04

Two assumptions carry the derivation: α and β stay the same every round, and distances d (not KLs) add up across rounds, which is how KL appears to grow in their runs. It can hold only up to some number of rounds, since the law is expected to fail at very short distances.

Why it matters

It puts a number on standard advice: refresh the reward model where the policy has gone. The gain grows only like log k, so the first few refreshes are worth the most, and none of them helps with the noise part.

4.4 Policy size independence · original

Everyday picture

If tuning a model is like updating a belief, a bigger model is a better-informed starting belief about what people write, not a more cunning optimizer.

Tiny example

Scaling the proxy from 12M upwards moves the peak out (§3.2). Scaling the policy from 1.2B to 6B does not move it, and the 6B run even travels a shorter KL for the same number of RL steps.

Why it matters

The paper's hypothesis, citing Korbak et al. (2022): RL with a KL penalty can be viewed as Bayesian inference starting from the initial policy as a prior, so a bigger policy mainly models the human demonstrations better. It is offered as a possibility, from two data points.

4.5 Limitations and future work · original

Everyday picture

This experiment studies a trainee copying a head judge. It cannot see whether the head judge was right.

Tiny example

A labeller may pick the answer that only looks right. The paper's footnote example is a robotic hand, trained from human feedback, that learned only to appear to grasp a ball (Christiano et al., 2017). That gap, between labels and intent, is a second source of overoptimization the synthetic setup leaves out by design. Towards Understanding Sycophancy studies one form of it: labels that favour answers agreeing with the person asking.

What the paper lists

  • Validate on other environments (they saw similar-looking results on WebGPT) and check that the synthetic setting transfers, since reward models may err in correlated ways.
  • Make reward models more robust to optimization; model the proxy score; study other optimizers (steering, beam search, other RL algorithms).
  • Explore adversarial Goodhart, more policy sizes, and multi-round RLHF.

Why it matters

A careful paper says what it did not measure. Here that is the biggest open question: how good are the labels themselves?

5 Related work · original

Everyday picture

The paper sits where four streams meet: Goodhart's law, reward hacking, adversarial robustness, and scaling laws.

The threads

  • Goodhart and reward hacking. Overoptimizing a reward model is a special case of specification gaming, also called reward hacking, with long catalogues of examples. Zhuang and Hadfield-Menell (2020) found very similar curves in toy environments.
  • Overfitting is Goodhart with a finite sample as the proxy.
  • Scaling laws (Kaplan et al., 2020, see the scaling laws companion) supplied the method: smooth trends that let small runs predict large ones.
  • RLHF from Christiano et al. (2017) to InstructGPT; Bai et al. (2022) also saw proxy scores roughly linear in √KL, though here there is a bend at small KL.

Why it matters

The contribution is not the observation that proxies fail (that was known) but the measurement of how, in a form that predicts.

Appendix B: other functional forms · original

Everyday picture

Before settling on a formula, you try its neighbours. The chosen one wins on predicting new data, not on elegance.

Tiny example

Two alternatives with a finite slope α at d = 0: d(α − β log(1 + d)), which extrapolated worse, and the power law d(α − β dγ), which needs a third number and whose best fits used small γ. Those small-γ fits are close to the chosen log form, because of a classic limit: for x = 5, 1000 × (51/1000 − 1) = 1.611, against ln 5 = 1.609.

In words: “as m grows, m times (the m-th root of x minus one) approaches the natural log of x.” So a power dγ with small γ behaves like 1 + γ log d, and the power law turns into the log law.

With the numbers: x = 5: m = 10 gives 1.746, m = 1,000 gives 1.611, and ln 5 = 1.609.

In Python:

import math
x = 5
# m (x^(1/m) − 1) for growing m, against log x
[round(m * (x ** (1 / m) - 1), 3) for m in (10, 1000)], round(math.log(x), 3)  # → ([1.746, 1.611], 1.609)

Why it matters

The log form is a simpler member of the power-law family, chosen because it predicted better with fewer knobs. (The paper's own limit uses n; this page writes m to keep n for the number of samples.)

Where it connects

In the paperWhere it connectsLearn it
A large model labels data in place of peopleLightman et al. (2023) cite this paper for the same move when comparing process and outcome supervision cheaplyLet's Verify Step by Step companion
Proxy of human labels versus human intent (§4.5)Sharma et al. (2023) cite it and study preference models rewarding agreement over truthSycophancy companion
Gold rises then falls with KLThe standard picture of reward hacking under a learned reward, with a KL leash as one defencereinforcement learning
Rewards that cannot be overoptimizedRule-based, checkable rewards remove the learned proxy altogether where the task allowsDeepSeekMath companion
Best-of-n against a learned scorerAlso used to pick answers at inference time, where the same peak-then-decline risk appliesTraining Verifiers companion

Glossary

Every term with hover guidance on this page, in one place.