InstructGPT, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same.
- The three-step diagram in §3.1 is live: tap a step to see what the data looks like at that point.
Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. The training stages lesson builds SFT loss masking and a reward-model loss in NumPy. The follow-up that removed the reinforcement learning step is the DPO companion.
Abstract · original
“Making language models bigger does not inherently make them better at following a user’s intent.”Ouyang et al. (2022), Abstract
Everyday picture
A base model like GPT-3 learned by predicting the next word on web pages, so it is a brilliant mimic, not a helpful assistant. Ask it “Write a poem about the sea” and it might continue with a list of other writing prompts, because that is what such a line is often followed by online. This paper teaches the model what people actually want by showing it examples, and then asking people to grade its attempts.
What the paper claims
- People preferred the answers of a 1.3-billion-parameter InstructGPT model over those of the 175-billion-parameter GPT-3: over 100 times smaller, yet more useful.
- InstructGPT is more truthful and somewhat less toxic than GPT-3, with small losses on public benchmarks that a tweak to training largely fixes.
- The method: supervised fine-tuning on human demonstrations, then reinforcement learning from human feedback.
Why it matters today
This three-step recipe (demonstrate, compare, reinforce) is the template behind the chat assistants that followed. Even where the final step has changed, the idea of shaping a model with human preferences comes from here.
1 Introduction · original
Everyday picture
The authors call the next-word objective misaligned: “predict what comes next on a web page” is a different job from “follow the user's instructions helpfully and safely”. They borrow a three-word target from earlier work: a model should be helpful (solve the user's task), honest (not make things up) and harmless (not cause harm).
The main findings, in one list
- Labelers preferred 175B InstructGPT to 175B GPT-3 85 ± 3% of the time, and to GPT-3 given a carefully written few-shot prompt 71 ± 4% of the time.
- On the TruthfulQA benchmark, InstructGPT gave truthful and informative answers about twice as often as GPT-3.
- On tasks where the answer should only use the input (such as summarizing), it made things up 21% of the time against GPT-3's 41%.
- When asked to be respectful, it produced about 25% fewer toxic outputs. It did not reduce measured bias.
- Scores on some public benchmarks dropped (an “alignment tax”), which mixing in pretraining data largely repaired.
- It generalized to labelers who wrote no training data, and to instructions it rarely saw, such as code questions and other languages.
- It still makes simple mistakes.
2 Related work · original
Learning from human preferences had already worked for robots in simulation, then for text summarization and story continuation. Other groups trained models on large collections of public NLP tasks phrased as instructions (FLAN and T0). The paper also draws on work measuring the harms of language models. Its new step was applying human-feedback training to a broad range of real user instructions on a very large model.
3 Methods · original
3.1 The three steps · original
Everyday picture
Training a new employee: first you show them how good work looks (demonstrations). Then you grade several of their attempts against each other, and a supervisor learns your taste from those grades (a reward model). Finally the employee practises on their own while the supervisor gives feedback, and they keep improving (reinforcement learning).
Tap any box, top to bottom in each column. Tap a column's dashed frame for the big picture of that step.
Reading it: read each column top to bottom, then left to right. In step 1, people write model answers for sampled prompts, and GPT-3 is fine-tuned to imitate them. In step 2, the step-1 model writes several answers per prompt, people rank them, and a separate reward model learns to predict their rankings. In step 3, the model writes answers to new prompts, the reward model scores each one, and the PPO algorithm nudges the model towards higher scores. The long arrow along the bottom shows the SFT model feeding the later steps. The paper notes steps 2 and 3 can be repeated: collect rankings of the improved model's answers, retrain the reward model, update again.
3.2–3.3 Prompts and tasks · original
Everyday picture
To learn what users want, train on what users ask. Most prompts came from people using an early version of the model through OpenAI's web interface (who had been told their data might be used for training), filtered for personal information and capped at 200 per user. Labelers also wrote prompts to get started.
| Dataset | Training prompts | Used for |
|---|---|---|
| SFT | ~13,000 | demonstrations, step 1 |
| Reward model | ~33,000 | rankings, step 2 |
| PPO | ~31,000 | prompts only (no labels), step 3 |
By use case, the prompts were mostly open-ended: generation 45.6%, open question answering 12.4%, brainstorming 11.2%, chat 8.4%, rewriting 6.6%, summarization 4.2%, and smaller shares of classification, closed question answering and extraction. Over 96% were in English. Labelers were asked to infer what the person who wrote each prompt wanted, including implicit wishes such as truthfulness.
Why it matters
Only about 18% of real prompts were the classification and question-answering tasks that public benchmarks mostly measure. That mismatch explains a result in §4.1.
3.4 The labelers · original
About 40 contractors, chosen by a screening test for sensitivity to different groups' preferences and skill at spotting harmful outputs. During training, labelers were told to prioritize helpfulness; in final evaluations, truthfulness and harmlessness. Agreement was high for such open-ended work: training labelers agreed with each other 72.6 ± 1.5% of the time, and held-out labelers 77.3 ± 1.3%.
Why it matters
“Aligned to human preferences” here means aligned to these people, following these instructions. The paper is candid about this in §5.2.
3.5 Models: SFT, reward model and PPO · original
Step 1: supervised fine-tuning
GPT-3 trained on the labelers' demonstrations for 16 epochs. An interesting detail: validation loss started overfitting after 1 epoch, yet more epochs still improved the reward model's score and human ratings. A training metric and the real goal can disagree.
Step 2: the reward model
Everyday picture: a wine judge who can't describe the perfect wine but reliably says which of two glasses is better. The reward model takes a prompt and an answer and returns a single number; it only has to get the ordering right.
Tiny example: a labeler ranks K = 4 answers: D best, then C, then A, then B worst. That ranking contains every pair: D over C, D over A, D over B, C over A, C over B, A over B. That is 6 comparisons, which is “4 choose 2”. Ranking 9 answers gives 36.
Reading it: each chip is one comparison the reward model learns from, written “winner ≻ loser”. The answers are lettered in ranked order, best first. Slide K from 4 to 9: the comparisons grow as K × (K − 1) / 2, from 6 to 36, from a single ranking task. That is why ranking is an efficient use of a labeler's time. The readout scores all the pairs with illustrative reward-model scores (each answer's score falls by 0.5 per rank) and averages the loss exactly as the equation below does.
In words: “for every pair from a ranking, the winner's score should beat the loser's; penalize by minus the log of the sigmoid of the gap, and average over all the pairs.”
With the numbers: if the reward model scores the winner 1.5 and the loser 0.5, the pair's loss is −ln σ(1.0) = −ln 0.731 = 0.313. A ranking of 4 answers contributes 6 such terms, each weighted by 1/6.
In Python:
import math
# σ
sigma = lambda z: 1 / (1 + math.exp(-z))
# r_θ(x, y_w), r_θ(x, y_l)
r_w, r_l = 1.5, 0.5
round(sigma(r_w - r_l), 3) # → 0.731
# one pair's loss
round(-math.log(sigma(r_w - r_l)), 3) # → 0.313
# C(K, 2) pairs in a ranking of K = 4
math.comb(4, 2) # → 6
Two practical details. The paper feeds all C(K, 2) comparisons from one ranking together as a single batch element, because treating them as independent examples made the reward model overfit after one pass. And it used a 6-billion-parameter reward model, since a 175B one trained unstably.
Step 3: reinforcement learning with PPO
Everyday picture: the model practises, the judge scores, and the model is nudged towards higher scores. But a clever student can learn to please a judge without actually getting better (for example by padding answers with phrases the judge likes). So there is a leash: a penalty, measured with the KL divergence, for drifting too far from the SFT model. The paper also mixes in ordinary next-word training on the original pretraining text, so the model doesn't forget general skills; this variant is called PPO-ptx, and it is what “InstructGPT” refers to.
In words: “for answers the model writes, earn the reward model's score minus a penalty for how much more likely the model now makes that answer than the SFT model did; separately, keep the model good at predicting ordinary pretraining text, weighted by γ.”
With the numbers (illustrative values, not the paper's settings): an answer scores 2.0 from the reward model; the RL model gives it a log-probability 3.0 higher than the SFT model did; with β = 0.1 that answer contributes 2.0 − 0.1 × 3.0 = 1.7. Setting γ = 0 gives the plain “PPO” model; γ > 0 gives PPO-ptx.
In Python:
# r_θ(x, y), the reward model's score
r = 2.0
# log(π_RL(y | x) / π_SFT(y | x))
log_ratio = 3.0
# β
beta = 0.1
round(r - beta * log_ratio, 1) # → 1.7
The paper found mixing in pretraining data worked better than simply tightening the leash (raising β), which cost reward without fully fixing the regressions.
Why it matters
This objective, reward minus a KL leash, is the same one DPO later solved without reinforcement learning. The training stages lesson shows both.
3.6 What “aligned” means here · original
Evaluation of helpfulness was by labeler preference: how often one model's answer beat a baseline's (the 175B SFT model, a mid-pack reference). Honesty is hard to measure directly, since you can't see what a model “believes”, so the paper measured truthfulness: made-up information on tasks that should only use the input (hallucination), and the TruthfulQA benchmark. Harm was approximated by specific questions, such as “inappropriate for a customer assistant?” or “denigrates a protected class?”, plus toxicity and bias benchmarks.
4 Results · original
4.1 On real user prompts · original
Everyday picture
A taste test: labelers see two answers to the same prompt and pick the better one. A win rate is how often a model's answer is picked.
Reading it: each bar is how often 175B InstructGPT's answers were preferred in a head-to-head comparison against the named opponent, as reported in the paper's text, with its 95% confidence range. The dashed line at 50% would mean “no better than the opponent”. Every bar clears it by a wide margin: against raw GPT-3 people picked InstructGPT 85% of the time, and even against GPT-3 helped by a careful few-shot prompt, 71%. The FLAN and T0 models were GPT-3 fine-tuned on large collections of public NLP tasks; they lost too.
- Size isn't everything: outputs of the 1.3B InstructGPT were preferred to those of the 175B GPT-3.
- More dependable: InstructGPT followed explicit constraints (“answer in 2 paragraphs or fewer”) more often, attempted the right task more often, and was more often appropriate for a customer assistant.
- Not just pleasing its own trainers: held-out labelers who wrote no training data preferred it at about the same rate. Reward models trained on 4 groups of labelers predicted a fifth group's choices with 69.6 ± 0.9% accuracy, against 72.4 ± 0.4% on their own groups.
- Why public task collections lost: they are mostly classification and question answering, about 18% of what users ask, while open-ended generation and brainstorming are about 57%.
4.2 On public benchmarks · original
| Question | Finding |
|---|---|
| Truthfulness (TruthfulQA) | Truthful and informative about twice as often as GPT-3 |
| Making things up (closed-domain API tasks) | 21% of the time, against 41% for GPT-3 |
| Toxicity (RealToxicityPrompts), when asked to be respectful | About 25% fewer toxic outputs than GPT-3 |
| Toxicity, with no instruction | About the same as GPT-3 |
| Toxicity, when explicitly asked to be toxic | Much more toxic than GPT-3: it follows instructions, including bad ones |
| Bias (Winogender, CrowS-Pairs) | No improvement over GPT-3 |
| Standard NLP benchmarks | PPO lost ground on SQuAD, DROP, HellaSwag and translation; PPO-ptx recovered most of it |
The alignment tax
Everyday picture: training an employee to be polite shouldn't make them worse at arithmetic. An “alignment tax” is exactly that kind of loss: capability given up for alignment. The authors argue it matters beyond this paper, because a high tax would tempt people to use unaligned but more capable models. Mixing pretraining updates into PPO (the γ term in §3.5) paid most of it back.
4.3 Qualitative results · original
InstructGPT could answer questions about code and follow instructions in other languages, though it often replied in English, even though both were rare in its training data. It generalized the idea of “following instructions”. It also made telling mistakes:
- It accepted false premises, happily explaining why something untrue is true.
- It hedged too much, listing possibilities when the answer was clear. The authors suspect labelers rewarded epistemic humility, and the reward model picked that up.
- It struggled with several constraints at once, such as “list 10 movies from the 1930s set in France”.
Why it matters
Each of these failure modes traces back to the training signal. A reward model learns what labelers rewarded, including their quirks. That lesson carries straight into how evaluations are designed today.
5 Discussion · original
Cheap, compared with pretraining
The whole alignment process was a small fraction of GPT-3's training compute.
Reading it: the bars compare training compute in petaflop/s-days (one petaflop per second sustained for a day), on a log scale, as reported in the paper. Fine-tuning the 175B SFT model took 4.9 and the PPO-ptx stage 60, against 3,640 for pretraining GPT-3: roughly 60 times less compute for the RL stage, for a gain labelers valued more than a 100× increase in model size.
Who are we aligning to?
“This procedure aligns the behavior of GPT-3 to the stated preferences of a specific group of people (mostly our labelers and researchers), rather than any broader notion of ‘human values’; we discuss this further in Section 5.2.”Ouyang et al. (2022), §1
The model is shaped by the labelers (mostly English-speaking, in the United States or Southeast Asia), by the researchers who wrote their instructions, and by the customers whose prompts were used. Labelers disagreed with each other about 27% of the time. The authors call for broader participation and for ways to condition models on the values of different groups.
Limitations and open questions
The models are neither fully aligned nor fully safe: they can still produce toxic or biased text, make things up, and follow harmful instructions. Open questions include adversarial data collection to find worst-case behaviour, combining preferences with other signals, and deciding what the model should do when a user's request conflicts with avoiding harm.
What changed since
| In the paper | Common today | Lesson or companion |
|---|---|---|
| PPO against a separate reward model | Often replaced or complemented by direct methods such as DPO | DPO companion |
| Preferences from ~40 human labelers | Human labels supplemented by AI feedback guided by written principles | training stages |
| ~13k demonstrations | Same insight at larger scale: a modest amount of high-quality data goes a long way | training stages |
| Human win rates as the main metric | Human ratings plus model judges calibrated against them | LLM-as-a-judge companion |
Glossary
Every term with hover guidance on this page, in one place.