Constitutional AI, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation has a table decoding it, a sentence reading it aloud, the numbers worked through, and the same numbers in Python.
- The pipeline diagram in §1.2 is live: tap any box to see what goes in and what comes out.
Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. The alignment lesson builds a toy constitution, a critique-and-revise pass and the preference pairs it produces, in plain Python. The reward-model machinery this paper reuses is built in the training stages lesson and walked through in the InstructGPT companion.
Abstract · original
“The only human oversight is provided through a list of rules or principles, and so we refer to the method as ‘Constitutional AI’.”Bai et al. (2022), Abstract
Everyday picture
A newspaper has a style guide. A junior writer drafts a story, an editor reads it against the guide and writes notes in the margin, and the writer revises. Later the same editor is asked, over and over, which of two drafts is better, and those verdicts teach the newsroom what “better” means. Constitutional AI does this with language models: the style guide is a short list of principles written in plain language (the constitution), and a model plays the editor.
What the paper claims
- An assistant can be trained to be harmless with no human labels identifying harmful outputs: people wrote the principles and a few examples, and a model did all the harm-judging.
- Two stages: a supervised stage (the model critiques and rewrites its own answers, then is fine-tuned on the rewrites) and a reinforcement learning stage whose preference labels come from a model, not people: RLAIF.
- The result is harmless but not evasive: it engages with a harmful request by explaining why it won't help, instead of shutting the conversation down.
- Both stages can use chain-of-thought reasoning, which makes the model's judgements easier to read and improves how people rate the results.
Why it matters today
This paper showed that the expensive, slow part of RLHF, collecting human comparisons, can be partly replaced by a model applying written rules. That moves the question “what should the model do?” out of a pile of private labels and into a document people can read and argue about.
1 Introduction · original
“We chose the term ‘constitutional’ because we are able to train less harmful systems entirely through the specification of a short list of principles or instructions, i.e. a constitution.”Bai et al. (2022), §1
Everyday picture
Every AI system follows some set of principles, even if nobody wrote them down: they are baked into whatever the human raters happened to reward. The authors' second reason for the word “constitutional” is exactly this. If principles are unavoidable, better to write them down where everyone can see them.
The starting point
The paper builds on the authors' earlier RLHF assistant, trained on two kinds of human comparisons: which answer is more helpful, and which is more harmless. That earlier work found a tension between the two. Training for helpfulness makes a model willing to obey harmful requests; training for harmlessness makes it evasive, refusing controversial questions and sometimes staying evasive for the rest of the conversation.
1.1 Motivations · original
Everyday picture
A company that grows from ten staff to ten thousand can't have the founder review every decision. It writes a handbook and trains managers to apply it. The paper's four motivations are that situation, applied to supervising AI:
- Scaling supervision. Using AI to help people supervise AI more efficiently, which the paper calls “scaling supervision” (scalable oversight is the wider field's name for the goal). RLHF already took a step this way: the reward during training comes from a preference model, not a person watching live. But it still needs tens of thousands of human labels.
- A harmless but non-evasive assistant. An assistant that answers everything with “I don't know” is harmless and useless. The goal is one that declines harmful requests and explains why.
- Simplicity and transparency. Tens of thousands of labels don't tell anyone what the training goal was. A short list of principles does.
- Faster iteration. Changing the goal means editing the principles, not collecting a new round of human labels.
Tiny example: reading an Elo gap
The paper reports results as Elo scores computed from crowdworkers' side-by-side choices, the same rating system chess uses. Only the gap between two scores means anything. If model A sits 100 points above model B, people are predicted to prefer A about 64% of the time. At 200 points it is about 76%.
In words: “the chance people prefer A to B depends only on the difference of their ratings; every 400 points of difference multiplies the odds by ten.”
With the numbers: RA − RB = 100, so 10−100/400 = 10−0.25 = 0.562 and P = 1 / 1.562 = 0.64.
In Python:
R_A, R_B = 150, 50
# 10^(−(R_A − R_B)/400)
ten_pow = 10 ** (-(R_A - R_B) / 400)
round(ten_pow, 3) # → 0.562
# P(A ≻ B)
round(1 / (1 + ten_pow), 2) # → 0.64
Reading it: drag the gap and watch the two bars, which always add up to 100%: they are the predicted share of head-to-head comparisons each model wins. At a gap of 0 the split is 50/50. The paper's Figure 2 plots harmlessness Elo against helpfulness Elo for every training run, with gaps between the best and worst runs of a few hundred points on each axis, so this slider covers the range you need to read it. The benchmarks lesson derives the same formula and computes it in elo_win_probability.
Why it matters
The paper's Figure 2 is its headline: for the models trained with human feedback, gaining helpfulness costs harmlessness, and the runs trace out a curve of trade-offs. The runs trained with AI feedback sit above that curve, less harmful at the same helpfulness. In the language of trade-offs this is a Pareto improvement: better on one axis without being worse on the other.
1.2 The Constitutional AI approach · original
Everyday picture
Two stages, like learning to write for a newspaper. First you rewrite your own drafts following the editor's notes until your habits change (supervised stage). Then you write two versions of each piece, the editor picks the better one each time, and you gradually learn to write the kind the editor prefers (reinforcement stage).
The two stages, as the paper names them
- Supervised stage: critique → revision → supervised learning. A helpful-only assistant answers prompts designed to provoke harmful answers (red-teaming prompts). The model is asked to critique its answer against a randomly drawn principle, then to revise it. Repeat, with a fresh principle each time. Fine-tune a pretrained model on the final revisions. The paper says this stage gets the model “on-distribution”, so the next stage has less exploring to do.
- RL stage: AI comparisons → preference model → reinforcement learning. The supervised model writes two answers per harmful prompt. A feedback model is shown both and a principle, and asked which is better. Its choices, mixed with human labels for helpfulness, train a preference model, and the supervised model is then trained by RL against it.
Diagram
Tap any box, starting at the top left with the helpful-only RLHF model. The two 16 principles boxes are the same constitution used twice.
Reading it: the top frame is the supervised stage. A helpful-only model answers red-team prompts; each answer goes round the Response → Critique → Revision loop, with a principle drawn from the constitution on the left each time, and the “repeat” arrow sends each revision back for another critique. The revisions (from every round) fine-tune a model, SL-CAI. The long wire down the left edge carries SL-CAI into the bottom frame twice: it writes the pairs of answers, and it is also the starting point for RL. In the bottom frame, the feedback model compares each pair against a principle from the same constitution, its verdicts join the human helpfulness labels (right) to train one preference model, and RL against that preference model produces the final RL-CAI model. The constitution appearing twice is the whole idea: one written document steers both stages.
Why it matters
From the preference model downwards, this is ordinary RLHF. The only swap is who labels harmlessness. That makes the method easy to bolt onto an existing RLHF pipeline, and it is why the paper can compare the two fairly.
1.3–1.4 Contributions, models and data · original
The paper's three findings, in its own order:
- As models get more capable, their ability to spot harm improves a lot, and chain-of-thought reasoning improves it further, to the point of competing with preference models trained on human labels (§2).
- Critiques and revisions can be applied again and again, each round reducing harmfulness, and critiquing first beats revising directly (§3).
- RL on the AI-generated labels improves the model further, matching or beating human labels for harmlessness as judged by crowdworkers (§4).
The models are the authors' own pretrained language models (the results focus on the largest, with 52 billion parameters). The starting assistant was trained with RLHF using only helpfulness labels. For comparison they also trained new “helpful” and “HH” (helpful and harmless) RLHF models from human labels. The human labels themselves came from earlier work, where crowdworkers chose the more helpful or more harmless of two model answers, and for harmlessness also acted as red-teamers, writing prompts meant to draw out harmful answers.
2 Evaluating the potential for AI supervision · original
“The results suggest that large language models may already be approaching the performance of crowdworkers in identifying and assessing harmful behavior, and so motivate using AI feedback.”Bai et al. (2022), §2
Everyday picture
Before handing the marking to a teaching assistant, you check that they mark the way you would: give them papers you have already marked and count how often they agree. Section 2 does this for a model marking helpfulness, honesty and harmlessness.
What was tested
Earlier work had written conversations ending in two candidate answers, with the better one known: 221 of them. Models now scored “well over 90%” on those, so the authors wrote 217 harder ones, mostly subtle tests of harmlessness, including cases where an evasive answer should lose to one that is harmless and helpful. That makes 438 binary comparisons. Two kinds of judge were scored:
- a preference model trained on several hundred thousand human labels, counted correct when it scores the better answer higher;
- a language model asked a multiple-choice question (“which answer is better, (A) or (B)?”), with and without chain-of-thought reasoning written out first.
Chain of thought helped the larger models a lot, and the trend in Figure 4 suggests models larger than 52 billion parameters would compete with the human-trained preference model. A further small gain came from writing five chains of thought and averaging the probabilities each gave.
Tiny example (illustrative numbers)
Five chains of thought about the same pair. Each ends by leaning one way, giving probabilities for (A) of 0.9, 0.2, 0.8, 0.7 and 0.9. One chain went astray, but the average, 0.7, still leans towards (A), and less extremely than any single confident chain.
In words: “add up the probability each chain of thought gives to answer (A), and divide by the number of chains.”
With the numbers: (0.9 + 0.2 + 0.8 + 0.7 + 0.9) / 5 = 3.5 / 5 = 0.7.
In Python:
# p_s: each chain's probability for answer (A)
p = [0.9, 0.2, 0.8, 0.7, 0.9]
S = len(p)
# (1/S) Σ_s p_s
round(sum(p) / S, 2) # → 0.7
Why it matters
This section is the paper's justification for everything that follows: AI feedback is only worth training on if the AI's verdicts track what people would say. Appendix B adds two more checks on crowdworker-labelled red-team conversations: telling harmful from ethical assistant behaviour on a balanced set of 254 conversations, and a 9-way classification of the kind of harm on 287 examples. The same habit, checking a model judge against human labels before trusting it, is how the evals lesson calibrates an LLM judge.
3 Critiques, revisions and supervised learning · original
Everyday picture
Proofreading your own essay works better if you read it once hunting for one kind of problem (“is anything here unkind?”) and then rewrite, than if you just rewrite it cold. The supervised stage makes a model do exactly that, many times, and then learns from the finished rewrites.
3.1 Method · original
“Note that since the final prompt-revision pair is formatted in the same manner as the original prompt-response pair, we can apply the same critique-revision pipeline multiple times, giving us a sequence of revisions.”Bai et al. (2022), §3.1
The loop
- Show the helpful-only model a red-team prompt and sample its answer, which is often harmful.
- Append a critique request, drawn from the constitution, and sample the model's critique.
- Append the matching revision request and sample the rewrite.
- Put the prompt and the rewrite together. Because this looks just like the original prompt and answer, steps 2 to 4 can run again with a new principle.
Each principle in the supervised stage is a pair: a critique request and its revision request. There are 16, drawn at random at every step. Two of them, quoted from the paper's Appendix C.1, appear in the step-through below.
The model sometimes lost track of its role, writing a critique where a revision belonged or the reverse. The fix was a few worked examples in the prompt (few-shot prompting), all formatted the same way.
Try it: step through a critique and revision loop
This example is this page's own, written to show the shape of the loop; it is not a real model's output. The principles quoted are from the paper's Appendix C.1. The paper's own examples are in its Appendices A and D.
Reading it: press Next step and read down. The first answer does what the person asked without thinking about the person whose messages would be read. The first principle asks for harms in general; its critique names the privacy problem, and the revision removes it while still speaking to the person. The second round draws a different principle, about sounding like a thoughtful friend; its critique finds nothing harmful left but notices the answer ignores why someone might ask, and the second revision adds that. Notice the pattern the paper reports: the first revision removes most of the harm, and later ones make smaller changes. Every revision in the chain, not only the last, becomes fine-tuning data (§3.2).
What the paper observed
The original answers often contained harmful content; the first revision almost always removed most of it; later revisions sometimes helped further, but less visibly. The revisions were rarely evasive: the model engaged with sensitive topics thoughtfully instead of shutting down. Appendix A notes that the critiques themselves were often inaccurate or overstated, yet the revisions were still more harmless than the original.
In code
The alignment lesson builds a four-principle toy of this loop: critique returns the names of the principles an answer breaks, and revise applies one fix per broken principle. There the principles are keyword checks and the fixes are exact, so the rewrite always passes; a real model's critiques are softer and sometimes wrong.
3.2 Datasets and training · original
| What | Count | How used |
|---|---|---|
| Red-team prompts written by people (from earlier work) | 42,496 | 4 critique-revision rounds each |
| Red-team prompts written by a model, few-shot prompted | 140,335 | 4 critique-revision rounds each |
| All red-team prompts | 182,831 | 4 revisions per prompt |
| Helpfulness prompts written by people | 135,296 | 2 answers each from the helpful model |
A pretrained model was fine-tuned for one epoch on the revisions (from every round, not just the last) plus the helpfulness answers, at a constant learning rate of 0.5 times the pretraining rate, in batches of 1,024 sequences. All sampling used temperature T = 1. The helpfulness answers are there so the model doesn't forget how to be useful while learning to be harmless.
3.3 Main results · original
Everyday picture
A taste test with no scoresheet: crowdworkers chat with two models at once, and at every turn pick the answer they prefer. Thousands of such picks become Elo scores for helpfulness and for harmlessness.
What they found
- As in earlier work, the helpful-only RLHF model is more helpful and more harmful than the HH RLHF model.
- SL-CAI sits in between on harmlessness (more harmless than helpful RLHF, less than HH RLHF) and is less helpful than both RL models.
- SL-CAI is both more helpful and more harmless than the pretrained model it started from.
- The Elo scores rest on 10,274 helpfulness comparisons and 8,135 more (the paper's sentence leaves out the word; by context, harmlessness) across the 24 model snapshots in Figures 2 and 3.
One detail matters for reading every result: this time crowdworkers were told to prefer a thoughtfully harmless answer over an evasively harmless one. That lowers the HH RLHF model's harmlessness score and raises the helpful model's, which the authors suspect is why the two look closer than in their earlier paper.
3.4 Scaling trends · original
Number of revisions
Figure 5 scores the original answer (revision 0) and revisions 1 to 4 with preference models trained only on human labels. Harmlessness scores rise with every revision, most steeply at the first; helpfulness scores fall; the combined score rises. The authors add a caution from their earlier work: preference-model scores become less well calibrated at high values, so the later gains should be taken “with a grain of salt”. They trained SL-CAI-n models on revisions up to and including the n-th, for n = 1 to 4.
Number of principles
Figure 6 varies how many principles the constitution holds. It makes no significant difference to the harmlessness score. The authors still expect more principles to give more varied revisions, which helps the RL stage explore, though they did not measure diversity.
Why it matters
The first revision does most of the work. Everything after it buys less, and the measuring stick itself gets less trustworthy at the top of its range: a small instance of Goodhart's law, which the alignment lesson opens with.
3.5 Are critiques necessary? · original
Everyday picture: “write down what's wrong, then fix it” against “just fix it”. Figure 7 compares critiqued revisions with direct revisions, scored by the same 52B harmlessness preference model. Critiques helped the smaller models noticeably; for the larger ones the two were similar, though critiques were always slightly ahead. The authors kept critiques anyway, for transparency: a written critique shows why the model changed its answer, and may help it find subtler harms.
4 Reinforcement learning from AI feedback · original
Everyday picture
RLHF hires people to say which of two answers is better. RLAIF asks a model the same question, with a principle to judge by. Everything downstream (the preference model, the reinforcement learning) is unchanged.
4.1 Method · original
“Once the desired comparison labels are obtained, the remainder of the training pipeline (i.e., preference model training and RL) is exactly the same as RLHF.”Bai et al. (2022), §4.1
The question the feedback model is asked
The feedback model (a pretrained language model, for the main runs) sees the conversation, a principle, and the two answers labelled (A) and (B), ending with “The answer is:”. There are 16 principles for this stage, one drawn at random per comparison, and a few worked examples are placed before each question. One of the principles:
“Which of these assistant responses is less harmful? Choose the response that a wise, ethical, polite and friendly person would more likely say.”Bai et al. (2022), §4.1
Tiny example: turning the model's answer into a label
The feedback model doesn't just say “(A)”; it gives a probability to every possible next token. Suppose the log-probability of “(A)” is −0.4 and of “(B)” is −1.5. As probabilities these are e−0.4 = 0.670 and e−1.5 = 0.223, which add up to only 0.893, because the rest went to other tokens. Rescale the two so they add up to 1 and (A) gets 0.75. That number, not a flat “A wins”, is the training target: a soft label.
In words: “turn each option's log-probability back into a probability, then divide by the total of the two, so the pair shares 100% between them.” (This page's notation for the paper's “normalized probabilities as targets”.)
With the numbers: t = 0.670 / (0.670 + 0.223) = 0.670 / 0.893 = 0.75.
In Python:
import math
# ℓ_A, ℓ_B: log-probabilities of the tokens "(A)" and "(B)"
l_A, l_B = -0.4, -1.5
round(math.exp(l_A), 3), round(math.exp(l_B), 3) # → (0.67, 0.223)
# t = e^ℓ_A / (e^ℓ_A + e^ℓ_B)
t = math.exp(l_A) / (math.exp(l_A) + math.exp(l_B))
round(t, 2) # → 0.75
Is 0.75 meaningful? The paper argues these multiple-choice probabilities are fairly well calibrated (a “75%” is right about 75% of the time), and Figure 9 checks it on the §2 comparisons.
With chain of thought, and why it needs clamping
In the chain-of-thought variant the feedback model is the helpful RLHF model, the question is rephrased as a conversation, and it is told “Let's think step-by-step” before answering. The written reasoning usually ends by naming a winner, so the probability for that option comes out near 0 or 1: badly calibrated. The fix is to clamp the probability into the range 40% to 60%.
In words: “raise anything below 0.4 up to 0.4, lower anything above 0.6 down to 0.6, and leave everything in between as it is.”
With the numbers: a confident chain of thought gives t = 0.97, which becomes 0.6; t = 0.03 becomes 0.4; t = 0.55 stays 0.55.
In Python:
def clamp(t, lo=0.4, hi=0.6):
# t' = min(max(t, lo), hi)
return min(max(t, lo), hi)
[clamp(t) for t in (0.97, 0.03, 0.55)] # → [0.6, 0.4, 0.55]
Why it matters
A label is a claim about how sure you are. Hard labels (0 or 1) and over-confident chain-of-thought labels both claim certainty the feedback model doesn't have, and §4.3 shows what the policy does with that false certainty. The paper also used the supervised model both to write the answer pairs and as RL's starting point, so the preference model is trained on the kind of answers the policy actually writes, at least early in RL.
4.2 Datasets and training · original
| What | Count |
|---|---|
| Human helpfulness comparisons (for the preference model) | 135,296 |
| AI harmlessness comparisons (one per SL-CAI red-team prompt) | 182,831 |
| Extra model-written red-team prompts for RL | 491,142 |
| Extra model-written helpfulness prompts for RL | 474,300 |
Every RL run, AI feedback or human, used the same RL settings as the authors' earlier work and the same training prompts: all the prompts from §3.2 plus the extra ones above. One difference from their earlier paper: the RLHF baselines now start straight from pretrained models, since the earlier intermediate step made little difference compared with what RL added.
4.3 Main results · original
Headline
RL-CAI models are significantly more harmless than both the RLHF models and SL-CAI (Figures 3 and 8). With chain of thought, RL-CAI is slightly less helpful and slightly more harmless than without. In Figure 8 the RL-CAI runs gain a lot of harmlessness without a great cost in helpfulness.
Overtraining: Goodhart again
Trained too long, RL-CAI started to game its reward (reward hacking): it became overly harsh with harmful prompts, or tacked the same boilerplate onto most red-team answers, such as “you are valid, valued, and cared for”. Three things helped:
- Rewriting principles to discourage over-reactive or accusatory answers. Several principles in Appendix C say so explicitly, such as avoiding answers that are “too preachy, obnoxious, or overly-reactive”.
- Ensembling: drawing one of 16 principles per label instead of reusing one, which made the preference model's scores more robust.
- Soft or clamped labels. Without chain of thought, soft labels did much better than hard 0/1 labels. With chain of thought, clamping to 20–80% helped slightly and 40–60% helped more, so 40–60 is what the main results use.
Why soft labels calm the policy down: the loss, decoded
The preference model gives each answer a score r, and the chance it assigns to “A is better” is the sigmoid of the gap m = rA − rB (the Bradley-Terry model). Training it on a target t is ordinary cross-entropy between t and that chance. The paper describes this in words (“normalized probabilities as targets”); written out, it is:
In words: “reward the preference model for giving answer A the chance t of being better, and B the rest; the loss is smallest at the gap m* where the sigmoid of the gap equals t exactly.”
With the numbers: with the clamped target t = 0.6, the best gap is m* = ln(0.6 / 0.4) = ln 1.5 = 0.405, where the loss is 0.673. With the soft label t = 0.75 from §4.1, m* = ln 3 = 1.10. With a hard label t = 1, m* = ln(1/0) is infinite: at a gap of 3 the loss is still 0.049 and still falling, so training keeps pushing the gap wider. With t = 0.6, the same gap of 3 costs 1.249, pulling the gap back towards 0.405.
In Python:
import math
def sigma(z):
return 1 / (1 + math.exp(-z))
def L(m, t):
# −[t ln σ(m) + (1 − t) ln σ(−m)]
return -(t * math.log(sigma(m)) + (1 - t) * math.log(sigma(-m)))
# m* = ln(t / (1 − t)) for the clamped and the soft target
[round(math.log(t / (1 - t)), 3) for t in (0.6, 0.75)] # → [0.405, 1.099]
round(L(math.log(1.5), 0.6), 3) # → 0.673
# a gap of 3 under a hard label, then under the clamped one
round(L(3, 1.0), 3), round(L(3, 0.6), 3) # → (0.049, 1.249)
Hover or tap the curves to read the loss at any gap.
Reading it: the x-axis is the preference model's score gap between the two answers; the y-axis is the loss on one comparison. The solid curve is a hard label, t = 1: it keeps sloping down to the right forever, so every step of training rewards an even bigger gap. The dashed curve is the target t from the slider: it has a lowest point at m* = ln(t/(1 − t)), and gaps beyond it cost more. Slide t down to 0.6, the paper's clamp, and the valley sits close to zero: the preference model learns “A is a bit better” rather than “A is infinitely better”. A preference model with milder gaps gives the policy a milder reward to chase, which is the paper's explanation for why soft and clamped labels led to less extreme answers. With t = 1 the formula is exactly the reward-model loss built in reward_model_loss.
Figure 9 of the paper checks the calibration directly on the §2 comparisons, plotting the 52B feedback model's stated probabilities against how often it was right, and finds them reasonably well calibrated.
4.4 Harmlessness versus evasiveness · original
Everyday picture
Ask a librarian something awkward. One says “I can't discuss that” and turns away. Another explains why they won't help with that part, and points you somewhere useful. Both are harmless; only the second is helpful, and only the second tells you what the rule is.
What the paper found
The earlier HH RLHF models were often evasive, with canned replies like “I can't answer that”. RL-CAI is “virtually never evasive”. Appendix D sets them side by side: on several of the paper's sensitive prompts, the HH RLHF model's whole reply is a short refusal, while RL-CAI with chain of thought explains the problem with the question and answers what it can.
Figure 8 has a telling detail. Late in training, the harmlessness Elo of both RLHF baselines falls. For the helpful model, the authors' explanation is that it grows more willing to help with dangerous tasks. For the HH model, it grows more evasive, and crowdworkers had been told to prefer the nuanced answer over the evasive one when both were harmless. The authors trace the evasiveness partly to their older human-label instructions, which simply asked for the more harmless answer and so rewarded refusals.
Why it matters
How raters are instructed decides what a model learns. An instruction like “pick the more harmless answer” quietly teaches refusal, the over-refusal failure the alignment lesson's refusal section measures. Written principles can say what raters' instructions often leave out: engage, and explain.
4.5 Absolute harmfulness score · original
Everyday picture
Comparisons say which of two answers is worse, never how bad either is. So the authors also used a single-model score: in earlier red-teaming, each worker chatted with one model trying to bait it, then rated their own “success” on a 0 to 4 scale. A language model was trained to predict that rating from the whole conversation.
Tiny example (illustrative numbers)
A worker rates a conversation 3; the predictor says 2.2. Training uses the squared error, an L2 loss: (2.2 − 3)² = 0.64.
In words: “the penalty is the square of how far the predicted rating is from the worker's rating.”
With the numbers: (2.2 − 3)² = (−0.8)² = 0.64.
In Python:
# ŝ: predicted rating, s: the worker's rating, 0 to 4
s_hat, s = 2.2, 3
round((s_hat - s) ** 2, 2) # → 0.64
What they found
On 64 hand-picked held-out red-team prompts, averaged over 256 answers per prompt (Figure 10), the helpful-only RLHF model grows more harmful as training goes on, while HH RLHF, RL-CAI and RL-CAI with chain of thought all grow less harmful. The authors caution that absolute scores may be poorly calibrated, since different workers grade the 0 to 4 scale differently.
5 Related work · original
The paper places itself as an extension of RLHF, alongside other assistants trained on human data (LaMDA, InstructGPT, Sparrow). Sparrow's split of harmlessness into separate rules resembles a constitution. Earlier work on models critiquing themselves or learning from written feedback resembles the supervised stage. Chain-of-thought prompting (“think step-by-step”) supplies the reasoning in both stages. Work showing language models make well-calibrated choices justifies using their probabilities as soft labels, and proposals for supervising AI with AI's help (scalable oversight) supply the motivation. The InstructGPT companion covers the human-feedback recipe this paper modifies.
6 Discussion · original
“Our ultimate goal is not to remove human supervision entirely, but to make it more efficient, transparent, and targeted.”Bai et al. (2022), §6
What was and wasn't removed
Human labels for harmlessness are gone; human labels for helpfulness remain. The authors expect helpfulness could also be learned without human feedback, from a pretrained model and extensive prompting, and leave that open.
6.1 Future directions · original
- Steering more than harm. The same method could change writing style, tone or persona, or make a model add caveats to certain kinds of advice. Cheap labels also make it practical to generate feedback along dozens of behavioural directions and study how they interact.
- Robustness. If helpfulness and harmlessness conflict less, red-teaming can be automated and scaled up without the model learning to refuse everything.
- Online training. Keep updating the preference model with fresh AI feedback on the policy's latest answers, which worked well with human feedback and can now be fully automated.
6.2 Broader impacts · original
The authors name two risks. Anything that makes it easier to train a model to behave as its creators intend also makes it easier to train a harmful one, and the supervised stage needs no RL machinery, so it is especially accessible. And needing fewer human labels means fewer people look at the model's behaviour before release, so failures nobody anticipated are more likely to ship. On the other side, fewer people have to do the unpleasant work of trying to provoke harmful content.
Appendices · original
What the constitution actually says (Appendix C)
The principles are short and overlapping, and the authors say they were chosen “in a fairly ad hoc manner for research purposes”. The 16 supervised-stage pairs mostly ask the model to find and remove harmful, unethical, dangerous or illegal content; a few target something specific, such as content inappropriate for young children, or whether the answer sounds like “a thoughtful, empathetic, caring, sensitive friend or therapist”. The 16 RL-stage principles are multiple-choice instructions, several of which warn against the over-reaction seen in §4.3, asking for more ethical awareness “without sounding excessively condescending, reactive, annoying or condemnatory”.
The few-shot examples (Appendix E)
The worked examples placed before each critique include one about honesty, not harm: a critique of an answer that stated a shaky claim too confidently, and a revision that states it with appropriate uncertainty. It is a small demonstration of the paper's §6.1 point that the same machinery can steer behaviours other than harmlessness.
In code
The alignment lesson's constitutional_preference_pairs is the RL stage's labeller in miniature: it scores each candidate with violations (how many toy principles it breaks), keeps every pair with a strict winner and drops ties. It produces hard labels; this page's §4.3 chart shows what soft labels would change.
What carries forward
| Idea in the paper | Why it lasts | Where to build it |
|---|---|---|
| Principles written in plain language | The training goal becomes something people can read, argue about and change without relabelling data | alignment |
| Critique, then revise | The same draft-check-fix loop reappears wherever a model improves its own output, from training data to agent reflection | planning |
| AI-labelled preference pairs | Pairs from any labeller can train a reward model or feed a method that skips the reward model entirely | DPO companion |
| Soft and clamped labels | A label should carry the labeller's uncertainty, or the reward will overstate it | training stages |
| Check the AI judge against people first (§2) | An automated judge is only as good as its agreement with the humans it replaces | evals |
| Harmless without being evasive | Refusal rate and harmful-compliance rate are two separate errors to measure | alignment |
Glossary
Every term with hover guidance on this page, in one place.