Direct Preference Optimization, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same. Hover the β in the DPO loss.
- The sliders in §4 compute the real DPO loss live. Drag them until the formula feels obvious.
Every idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. This paper builds on the pipeline introduced by InstructGPT; reading that companion first helps. The training stages lesson computes the DPO loss in NumPy.
Abstract · original
“In this paper we introduce a new parameterization of the reward model in RLHF that enables extraction of the corresponding optimal policy in closed form, allowing us to solve the standard RLHF problem with only a simple classification loss.”Rafailov et al. (2023), Abstract
Everyday picture
The standard way to teach a chat model manners, RLHF, is like training a dog through an interpreter. First you train a judge (a reward model) to score answers the way people would. Then the model practises answering, the judge scores each attempt, and a reinforcement learning algorithm slowly improves the model. It works, but it is slow, fiddly and unstable. DPO says: skip the judge. Show the model pairs of answers where people preferred one, and adjust it directly so it favours the preferred one, using an ordinary classification-style loss.
What the paper claims
- DPO optimizes the same objective as RLHF, exactly, without a separate reward model and without sampling from the model during training.
- It is stable and needs almost no hyperparameter tuning.
- It matched or beat PPO-based RLHF on controlling sentiment, summarization and single-turn dialogue, with models of up to 6 billion parameters.
Why it matters today
DPO and its many variants became a standard way to do preference tuning, especially for open models, because it is far simpler to run than a reinforcement learning loop.
1 Introduction · original
Everyday picture
A pretrained model has read everything: brilliant code and buggy code, facts and common misconceptions. The paper's example: we want the model to know a misconception that half of people believe, but not to assert it half of the time. Preference tuning selects the behaviour we want from everything the model can do.
Why not just use RLHF?
RLHF trains several models (the policy, a reward model, a value model) and samples fresh text from the policy inside the training loop. That is expensive and famously finicky. DPO's key observation is mathematical: the RLHF objective can be rewritten so that the language model itself is the reward model, and then a simple loss on preference pairs solves it.
2 Related work · original
Instruction tuning (training on instructions paired with good answers, called SFT) makes models follow instructions. But comparing two answers is easier for people than writing a perfect one, so later work trained on human preferences, first fitting a reward model and then optimizing it with reinforcement learning algorithms such as PPO. Related threads studied learning from preferences in bandit and reinforcement learning settings. DPO's contribution is a single-stage method that learns a policy straight from offline preference pairs.
3 Preliminaries: the RLHF pipeline · original
Everyday picture
RLHF has three phases. SFT: teach the format by example. Reward modelling: show people two answers to the same prompt, record which they preferred, and train a judge to predict that choice. RL fine-tuning: let the model write answers, score them with the judge, and push the model towards high scores without drifting too far from where it started.
Hover or tap a box. Compare the two rows.
Reading it: both rows start from the same data: prompts with a preferred and a rejected answer. The top row (RLHF) trains a reward model first, then runs a reinforcement learning loop in which the model writes new answers that the reward model scores, over and over (the feedback arrow). The bottom row (DPO) goes straight from preference pairs to the final model with one loss, using a frozen copy of the starting model (π_ref) as a reference point. Same destination, one box and one loop fewer.
The reward model: Bradley-Terry
How do you turn “people preferred A over B” into numbers? The Bradley-Terry model says: give every answer a hidden score, and the chance that A beats B depends only on the difference of their scores, passed through a sigmoid. It is how chess ratings work: a 200-point gap means a predictable win rate, whatever the absolute ratings.
In words: “the chance people prefer answer 1 to answer 2 is the sigmoid of how much higher answer 1's hidden score is.”
With the numbers: if answer 1 scores 1.5 and answer 2 scores 0.5, the gap is 1.0 and σ(1.0) = 1 / (1 + e−1) = 0.731: people pick answer 1 about 73% of the time.
In Python:
import math
# r*(x, y1), r*(x, y2): the hidden scores
r1, r2 = 1.5, 0.5
# the fraction form
round(math.exp(r1) / (math.exp(r1) + math.exp(r2)), 3) # → 0.731
# σ, the sigmoid
def sigma(t):
return 1 / (1 + math.exp(-t))
# the same number: σ of the gap
round(sigma(r1 - r2), 3) # → 0.731
Reading it: the x-axis is how much higher answer 1 scores than answer 2; the y-axis is the chance people prefer it. At a gap of 0 it is a coin flip (0.5). A gap of +2 gives about 88%, and the curve flattens towards 1 but never reaches it. Only the gap matters: adding 100 to both scores changes nothing. Remember that; it is the fact DPO's derivation relies on.
Training the reward model is then ordinary binary classification on “which one won”:
In words: “on average over the labelled pairs, make the winner's score beat the loser's by enough that the sigmoid of the gap is close to 1.”
With the numbers: winner scored 1.5 and loser 0.5, so the loss for that pair is −ln σ(1.0) = −ln 0.731 = 0.313. Widen the gap to 3 and the loss falls to −ln σ(3) = −ln 0.9526 = 0.049 (round σ(3) to 0.953 first and you get 0.048).
In Python:
import math
def sigma(t):
return 1 / (1 + math.exp(-t))
# r_φ(x, y_w), r_φ(x, y_l)
r_w, r_l = 1.5, 0.5
# −log σ(r_w − r_l) for this one pair
round(-math.log(sigma(r_w - r_l)), 3) # → 0.313
# widen the gap to 3
round(sigma(3), 3), round(-math.log(sigma(3)), 3) # → (0.953, 0.049)
The RL objective
With a reward model in hand, RLHF maximizes reward while staying close to the starting model. The closeness penalty is the KL divergence: a measure of how different two probability distributions are. Without it, the model would find nonsense that happens to fool the reward model.
In words: “find the model that earns the most reward on average, minus a penalty, scaled by β, for how far its answers drift from the reference model's.”
With the numbers: if a candidate model earns average reward 2.0 but has drifted by KL = 5 (in nats, the natural-log unit), then with β = 0.1 its score is 2.0 − 0.1 × 5 = 1.5. A model earning 1.8 with KL = 1 scores 1.7 and wins: less reward, much less drift.
In Python:
beta = 0.1
# E[r_φ(x, y)] − β · KL[π_θ ‖ π_ref]
def score(reward, kl):
return reward - beta * kl
round(score(2.0, 5), 2), round(score(1.8, 1), 2) # → (1.5, 1.7)
# max over the candidate models
max([(2.0, 5), (1.8, 1)], key=lambda m: score(*m)) # → (1.8, 1)
Why it matters
Optimizing this needs reinforcement learning because generated text is discrete: you can't take a gradient through “sample a word”. That, and the instability of PPO, is the pain DPO removes. See the InstructGPT companion for the pipeline in practice.
4 Direct Preference Optimization · original
This section is the heart of the paper: three short algebraic moves that turn the RL problem into a classification loss.
Step 1: the best possible model has a formula
Everyday picture
Imagine the reference model as a shop's current stock levels, and the reward as how much customers like each product. The best new stock under “don't change too much” is simple: take the current stock and boost each product in proportion to e(how liked)/β, then rescale so the shelves add up to 100%. That rescaling constant is Z.
In words: “the optimal model is the reference model reweighted by e to the reward over β, then normalized so the probabilities add to 1.”
With the numbers: two possible answers, each with reference probability 0.5; rewards 1 and 0; β = 1. Unnormalized weights: 0.5 × e1 = 1.359 and 0.5 × e0 = 0.5, so Z = 1.859. The optimal model gives the first answer 1.359 / 1.859 = 0.731 and the second 0.269. With β = 0.1 instead, the first answer's weight becomes 0.5 × e10, and the optimal model picks it almost always: small β means “chase reward”, large β means “stay close to the reference”.
In Python:
import math
pi_ref, r, beta = [0.5, 0.5], [1, 0], 1
# π_ref(y|x) · exp(r(x,y) / β)
weights = [p * math.exp(r_y / beta) for p, r_y in zip(pi_ref, r)]
[round(w, 3) for w in weights] # → [1.359, 0.5]
# Z(x): makes the shares add to 1
Z = sum(weights)
round(Z, 3) # → 1.859
# π_r(y|x)
[round(w / Z, 3) for w in weights] # → [0.731, 0.269]
# chase reward harder
beta = 0.1
weights = [p * math.exp(r_y / beta) for p, r_y in zip(pi_ref, r)]
# almost always the first answer
weights[0] / sum(weights) > 0.9999 # → True
The catch: Z(x) sums over every possible answer to the prompt, which is astronomically many. You can't compute it, so you can't use this formula directly.
Step 2: flip it around
Everyday picture
If you can't compute the best stock from the customer ratings, run it backwards: look at any stock level and infer what ratings would have made it the best. Taking logarithms of the formula above and rearranging gives the reward in terms of the model.
In words: “a reward is β times how much more likely the model makes an answer than the reference did (on a log scale), plus a term that depends only on the prompt.”
With the numbers: for the first answer, 1 × ln(0.731 / 0.5) + 1 × ln 1.859 = 0.380 + 0.620 = 1.000, the reward we started with. For the second, ln(0.269 / 0.5) + 0.620 = −0.620 + 0.620 = 0. The formula recovers both rewards exactly.
In Python:
import math
beta, pi_ref = 1, 0.5
# Z(x) from step 1
Z = 0.5 * math.e + 0.5
# the optimal model from step 1
pi_r = [0.5 * math.e / Z, 0.5 / Z]
round(beta * math.log(pi_r[0] / pi_ref), 3), round(beta * math.log(Z), 3) # → (0.38, 0.62)
# r = β log(π_r / π_ref) + β log Z
r_first = beta * math.log(pi_r[0] / pi_ref) + beta * math.log(Z)
r_second = beta * math.log(pi_r[1] / pi_ref) + beta * math.log(Z)
# + 0 tidies the float's -0.0 into 0.0
round(r_first, 3), round(r_second, 3) + 0 # → (1.0, 0.0)
The trick: Z cancels
Recall from Bradley-Terry that only the difference of two rewards matters. Both answers to the same prompt share the same β log Z(x), so in the difference it cancels. The uncomputable term vanishes. That is the whole paper in one sentence.
Step 3: the DPO loss · original
Plug the flipped reward into the reward-model loss from §3. What remains is a loss on the language model itself:
In words: “for each labelled pair, measure how much more the model now likes the winner than the reference did, subtract the same for the loser, scale by β, and push that margin up through the same log-sigmoid loss a reward model would use.”
With the numbers: β = 0.1. The model gives the winner a log-probability of −12.0 where the reference gave −14.0 (a gain of +2.0), and the loser −15.0 where the reference gave −13.0 (a loss of 2.0). The implicit rewards are 0.1 × 2.0 = 0.2 and 0.1 × (−2.0) = −0.2, so the margin is 0.4. σ(0.4) = 0.599 and the loss is −ln 0.599 = 0.513.
In Python:
import math
def sigma(t):
return 1 / (1 + math.exp(-t))
beta = 0.1
# log π_θ(y_w|x), log π_ref(y_w|x)
logp_w, logp_ref_w = -12.0, -14.0
# log π_θ(y_l|x), log π_ref(y_l|x)
logp_l, logp_ref_l = -15.0, -13.0
# β log(π_θ / π_ref) for the winner
r_w = beta * (logp_w - logp_ref_w)
# ... and for the loser
r_l = beta * (logp_l - logp_ref_l)
round(r_w, 2), round(r_l, 2), round(r_w - r_l, 2) # → (0.2, -0.2, 0.4)
# σ(margin), −log σ(margin)
round(sigma(r_w - r_l), 3), round(-math.log(sigma(r_w - r_l)), 3) # → (0.599, 0.513)
Try it: the DPO loss, live
Reading it: the first two sliders say how much the model being trained has raised (positive) or lowered (negative) the log-probability of the preferred and rejected answers, compared with the frozen reference. The bars show the two implicit rewards, then the loss and the size of the push the next update will get. Raise the preferred answer and lower the rejected one: the loss falls towards 0 and the push fades, because this pair is “learned”. Flip them, so the model prefers the loser: the loss climbs and the push grows. Now change β: it sets how much a given change in log-probability counts, so a larger β reaches a low loss with a smaller departure from the reference.
Hover the curves. The x-axis is the log-ratio margin: (winner's gain) − (loser's gain), before scaling by β.
Reading it: each curve is the DPO loss for one β, plotted against the unscaled margin. Right of zero, the model already prefers the winner more than the reference did, and the loss decays towards 0; left of zero it rises almost in a straight line. Larger β makes the curve steeper, so the same margin counts for more. At a margin of 4 with β = 0.1 the loss is still 0.513, the worked example above. The same margin gives 0.127 with β = 0.5 and 0.018 with β = 1.
What does the update actually do?
“Intuitively, the gradient of the loss function ℒDPO increases the likelihood of the preferred completions yw and decreases the likelihood of dispreferred completions yl.”Rafailov et al. (2023), §4
Everyday picture
A good tutor spends time on the questions you get wrong, not on the ones you have mastered. The DPO gradient has a built-in weight that does exactly this: pairs the model currently ranks the wrong way get a big push, and pairs it already ranks correctly get a small one.
In words: “move the weights to make the winner more likely and the loser less likely, weighted by how badly the model's implicit rewards currently rank this pair.”
With the numbers: in the worked example the implicit rewards were 0.2 (winner) and −0.2 (loser), so the weight is σ(−0.2 − 0.2) = σ(−0.4) = 0.401. Had the model ranked them the wrong way round (−0.2 and 0.2), the weight would be σ(0.4) = 0.599: a bigger push for the mistake.
In Python:
import math
def sigma(t):
return 1 / (1 + math.exp(-t))
# r̂_θ(x, y_w), r̂_θ(x, y_l)
r_w, r_l = 0.2, -0.2
# the weight σ(r̂_l − r̂_w)
round(sigma(r_l - r_w), 3) # → 0.401
# ranked the wrong way round
r_w, r_l = -0.2, 0.2
round(sigma(r_l - r_w), 3) # → 0.599
The paper notes that dropping this weight, so every pair is pushed equally, makes the model degenerate. The weight is not decoration.
The recipe
- Take prompts, have the reference model produce two answers each, and have people mark which is better. Or reuse an existing preference dataset.
- Freeze a copy of the starting model as πref (usually the SFT model).
- Minimize ℒDPO with ordinary gradient descent. In the paper: β = 0.1 by default (0.5 for summarization), batch 64, learning rate 10−6 with a 150-step warmup.
Why it matters
Everything in the recipe is supervised learning. No sampling during training, no reward model to keep in memory, no value network, no reinforcement learning tricks. Compute and memory drop, and so does the number of things that can go wrong.
5 Theoretical analysis · original
5.1 Your language model is secretly a reward model · original
Everyday picture
If two scoring systems always differ by the same amount on every answer to a given question (say, one adds 10 points to everything), they rank answers identically and lead to the same best model. The paper calls such rewards equivalent. It then proves that every class of equivalent rewards contains one that can be written as β log(π / πref) for some model π. So restricting ourselves to rewards of that form loses nothing.
In words: “any reward worth learning can be expressed as how much a model favours an answer relative to the reference, scaled by β.”
With the numbers: in the two-answer example, this form gives rewards 0.380 and −0.620 rather than 1 and 0. They differ from the originals by the same −0.620, so they rank the answers the same way and lead to the same optimal model.
In Python:
import math
beta, pi_ref = 1, 0.5
Z = 0.5 * math.e + 0.5
# the optimal model from step 1
pi = [0.5 * math.e / Z, 0.5 / Z]
# r = β log(π / π_ref)
r = [beta * math.log(p / pi_ref) for p in pi]
[round(r_y, 3) for r_y in r] # → [0.38, -0.62]
# the same shift for both answers
[round(new - old, 3) for new, old in zip(r, [1, 0])] # → [-0.62, -0.62]
Why it matters
This is the justification for the paper's subtitle: a language model trained with DPO carries a reward model inside it, readable as β log(π / πref). Practitioners use exactly that quantity to inspect what a DPO model has learned.
5.2 Why actor-critic RL is unstable · original
Viewed through the same lens, PPO-style RLHF implicitly has to estimate the normalizing term (the β log Z piece) from samples, usually with a learned value function or a rough baseline. Estimating it badly gives high-variance gradients and unstable training. DPO's reparameterization makes that term cancel exactly, so no baseline is needed.
6 Experiments · original
Everyday picture
Three tasks, rising in realism. Sentiment: continue movie-review openings so they come out positive, where a sentiment classifier provides the true reward, so the reward-versus-drift trade-off can be measured exactly. Summarization: summarize Reddit posts (TL;DR). Dialogue: answer single-turn requests from the Anthropic Helpful and Harmless dataset of 170,000 dialogues.
6.1 The reward versus drift frontier
On sentiment, across 22 runs with different settings, DPO reached the highest reward at every level of KL drift. It beat PPO even when PPO was given the true reward instead of a learned one. Both optimize the same objective; DPO simply optimizes it better.
6.2 Real preference data
- Summarization: judged by GPT-4 against human-written reference summaries, DPO won about 61% of the time at sampling temperature 0, against about 57% for PPO at its best temperature. DPO also held up better at higher temperatures.
- Head to head with humans judging: DPO summaries at temperature 0.25 were preferred 58% of the time over PPO's best.
- Dialogue: DPO was the only computationally efficient method that improved on the dataset's own preferred answers, matching an expensive best-of-128 sampling baseline.
6.3 A new input distribution
| Method | Temperature 0 | Temperature 0.25 |
|---|---|---|
| DPO | 0.36 | 0.31 |
| PPO | 0.26 | 0.23 |
Trained on Reddit and tested on news, DPO kept its lead: early evidence that it generalizes at least as well as PPO.
6.4 Can GPT-4 stand in for human judges?
| DPO | SFT | PPO-1 | |
|---|---|---|---|
| Number of respondents | 272 | 122 | 199 |
| GPT-4 (simple prompt) win % | 47 | 27 | 13 |
| GPT-4 (concise prompt) win % | 54 | 32 | 12 |
| Human win % | 58 | 43 | 17 |
| GPT-4 (simple) agrees with humans % | 70 | 77 | 86 |
| GPT-4 (concise) agrees with humans % | 67 | 79 | 85 |
| Humans agree with each other % | 65 | - | 87 |
GPT-4 agreed with humans about as often as humans agreed with each other. It also preferred longer, repetitive summaries until its prompt asked for conciseness, a reminder that LLM judges need to be checked against people. See the LLM-as-a-judge companion and the evaluation lesson.
7 Discussion · original
“With virtually no tuning of hyperparameters, DPO performs similarly or better than existing RLHF algorithms, including those based on PPO; DPO thus meaningfully reduces the barrier to training more language models from human preferences.”Rafailov et al. (2023), §7
The authors list open questions: how DPO generalizes compared with an explicit reward model; whether a DPO model can usefully label its own new data; how “reward over-optimization” (getting worse by optimizing too hard) shows up here; scaling beyond 6 billion parameters; and how much GPT-4 judgements depend on the prompt.
What changed since
| In the paper | Common today | Lesson |
|---|---|---|
| DPO on models up to 6B | DPO and variants applied to much larger open models | training stages |
| Human-labelled pairs | Pairs labelled by other models, guided by written principles, checked by people | evals |
| One β for everything | Many variants that change the loss shape or drop the reference model; the core idea is the same | losses |
| Offline pairs from the reference model | Often mixed with fresh pairs sampled from the model being trained | training stages |
Glossary
Every term with hover guidance on this page, in one place.