Classifier-Free Diffusion Guidance, annotated
How to read this page
Nothing on this page assumes you already know the jargon. Three things help:
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation is followed by a table of its symbols.
- The diagrams and pictures are live: hover or tap the parts, drag the sliders, and the readouts say what changed.
Each idea is explained in the same order: an everyday picture, a tiny example you could check by hand, a diagram, then the math, and finally why it still matters.
Two warnings about notation. First, this paper writes a noisy picture as z and the clean one as x, and it measures the noise level by λ, the log signal-to-noise ratio. Its αλ is a scale, not a share: z = αλx + σλε with αλ² + σλ² = 1. So αλ here is √ᾱt in the DDPM companion and √αt in the DDIM companion.
Second, the guidance weight w is counted from a different zero than in the diffusion lesson. Here w = 0 means “the plain labelled guess, no guidance”. In the lesson, w = 1 means that. So the lesson's weight is this paper's w plus 1: the lesson's w = 3 is the paper's w = 2.
The toys. The one-pixel toy from the DDIM companion: a clean value x = 2.0, noised to a level with αλ = 0.8 and σλ = 0.6 by the noise ε = 0.5, gives z = 1.9. And a two-class toy for guidance: class “+” is the clean value +2 and class “−” is −2, equally common. At the same noise level each class's noisy values form a bell curve centred at ±1.6 with variance 0.36, and we look at the ambiguous noisy value z = 0.1.
Abstract
“We show that guidance can be indeed performed by a pure generative model without such a classifier…”Ho and Salimans (2022), Abstract. Read the original
Everyday picture
A caricaturist draws a famous face by noticing what makes it different from an average face (the big chin, the heavy eyebrows) and exaggerating exactly that difference. Classifier-free guidance does the same: it asks a diffusion model “what would you draw if told the label?” and “what would you draw anyway?”, then pushes further in the direction of the difference.
What the paper claims
- Classifier guidance, an earlier method, trades sample diversity for sample quality after training, like the truncation trick of GANs, but needs a separately trained image classifier.
- The same trade-off can be had with no classifier: train one diffusion model both with and without its label, and at sampling time mix its two noise guesses.
- This classifier-free guidance reaches a quality/diversity trade-off similar to classifier guidance on ImageNet.
Why it matters today
Almost every text-to-image and text-to-video diffusion system uses this trick; the “guidance scale” slider in image tools is its w. It is also why each sampling step of such systems usually costs two network passes.
1 Introduction · original
“Classifier guidance complicates the diffusion model training pipeline because it requires training an extra classifier, and this classifier must be trained on noisy data so it is generally not possible to plug in a pre-trained classifier.”Ho and Salimans (2022), §1
Everyday picture
Imagine steering a sketch artist with an art critic who shouts “more dog!” after every stroke. It works, but you must hire a critic who can recognise a dog in a half-finished, smudged sketch, and ordinary critics trained on finished paintings cannot. Worse, an artist who learns to please one critic may learn to fool it rather than to draw better dogs.
Tiny example
An off-the-shelf classifier has only ever seen clean images like x = 2.0. During sampling it would be shown noisy values like z = 0.1 at a heavy noise level, where both classes are plausible (the two-class toy gives 71% “+”). A classifier for guidance must be trained on exactly such noisy inputs, at every noise level.
The paper's questions
- Can guidance be done without a classifier?
- Classifier guidance steps along a classifier's gradient, which looks like a gradient-based adversarial attack on that classifier. Does it score well on classifier-based metrics like Inception score and FID only because it fools classifiers?
- The answer to both, shown here: guidance works with a pure generative model, and reaches the same trade-off without any classifier gradient.
Why it matters
Removing the classifier removed a whole second model from the pipeline, and with it the limit that guidance only works for labels a classifier can recognise. That is what let guidance follow free text prompts.
2 Background · original
Everyday picture
Instead of numbering noise levels 1 to 1,000, the paper describes each level by a single dial: the log of how much louder the signal is than the noise. Turn the dial down and the picture fades into static; the model learns to turn it back up. This is the same forward process as DDPM, just labelled by a continuous number instead of a step count.
Tiny example
The toy pixel at z = 1.9 has signal variance 0.8² = 0.64 and noise variance 0.6² = 0.36. The ratio is 0.64/0.36 = 1.778, and its logarithm is λ = 0.575.
In words: “a noisy picture at level λ is the clean picture scaled by αλ plus noise of variance σλ²; the two variances always add up to 1, and λ = log(αλ²/σλ²) says how they are split.”
With the numbers: λ = log(0.64/0.36) = 0.5754; then 1/(1 + e−0.5754) = 1/(1 + 0.5625) = 0.64 = αλ², and σλ² = 0.36. The toy pixel: 0.8 × 2.0 + 0.6 × 0.5 = 1.9. At λ = 0 signal and noise are equal (αλ² = 0.5); the forward process runs toward smaller λ.
In Python:
import math
# λ: the log of signal variance over noise variance
lam = math.log(0.64 / 0.36)
round(lam, 4) # → 0.5754
# α_λ² = 1 / (1 + e^(−λ)) and σ_λ² = 1 − α_λ²
alpha2 = 1 / (1 + math.exp(-lam))
round(alpha2, 4), round(1 - alpha2, 4) # → (0.64, 0.36)
# z = α_λ x + σ_λ ε
round(math.sqrt(alpha2) * 2.0 + math.sqrt(1 - alpha2) * 0.5, 4) # → 1.9
Moving between two levels
Going from a cleaner level λ′ to a noisier level λ is itself a small forward step (equation 2), and running it backwards when the clean picture is known gives a bell curve with a known centre and width (equation 3), exactly the DDPM backward step of the DDPM companion:
In words: “the forward step between two levels adds a share (1 − eλ−λ′) of the noisier level's noise; run backwards knowing x, the cleaner level is a blend of the noisy value and the clean one, with a smaller spread.”
With the numbers: take the cleaner level λ′ = log(0.96/0.04) = 3.178. Then eλ−λ′ = 1.778/24 = 0.0741, so the forward step adds noise of variance 0.9259 × 0.36 = 0.3333. Backwards from z = 1.9 with x = 2.0: 0.0741 × (0.9798/0.8) × 1.9 + 0.9259 × 0.9798 × 2.0 = 1.9868, with variance 0.9259 × 0.04 = 0.0370. Both numbers match the DDIM companion's toy step exactly.
In Python:
import math
lam, lam_p = math.log(0.64 / 0.36), math.log(0.96 / 0.04)
a, a_p = 0.8, math.sqrt(0.96)
s2, s2_p = 0.36, 0.04
e = math.exp(lam - lam_p)
round(e, 4), round((1 - e) * s2, 4) # → (0.0741, 0.3333)
# μ̃ and σ̃² going back from z = 1.9 with x = 2.0
z, x = 1.9, 2.0
round(e * (a_p / a) * z + (1 - e) * a_p * x, 4), round((1 - e) * s2_p, 4) # → (1.9868, 0.037)
The reverse process and the training loss
The model's step backwards is the same bell curve with the model's guess of x in place of the real one. Its width is a blend, set by a knob v, of the two widths above:
In words: “guess the clean picture by removing the guessed noise, take the known backward step toward that guess, and pick a spread between the small backward one (v = 0) and the larger forward one (v = 1).”
With the numbers: with a perfect noise guess of 0.5, xθ = (1.9 − 0.6 × 0.5)/0.8 = 2.0. With v = 0.3 (the paper's 64 × 64 setting), the spread is 0.03700.7 × 0.33330.3 = 0.0716: about twice the small one, a fifth of the large one.
In Python:
z, sigma, alpha, eps_hat = 1.9, 0.6, 0.8, 0.5
# x_θ(z) = (z − σ ε_θ) / α
round((z - sigma * eps_hat) / alpha, 4) # → 2.0
# the reverse variance: (σ̃²)^(1 − v) (σ²_(λ|λ′))^v, with v = 0.3
v, small, large = 0.3, 0.037037, 0.333333
round(small ** (1 - v) * large ** v, 4) # → 0.0716
Training is DDPM's noise-guessing loss, with the level λ drawn at random:
In words: “noise a real picture to a random level, ask the network for the noise, and score the squared miss; pick the level by pushing a uniform number u through a curve that spends more time near the middle levels.”
With the numbers: one example with the guess 0.4 against the true 0.5 scores 0.01. With the paper's endpoints λmin = −20 and λmax = 20, u = 0 gives λ = 20 (almost clean), u = 0.5 gives λ = 0 (half signal, half noise), and u = 1 gives −20 (almost pure noise).
In Python:
import math
round((0.4 - 0.5) ** 2, 4) # → 0.01
lam_min, lam_max = -20, 20
b = math.atan(math.exp(-lam_max / 2))
a = math.atan(math.exp(-lam_min / 2)) - b
def lam(u):
# + 0.0 turns a rounded −0.0 into 0.0
return round(-2 * math.log(math.tan(a * u + b)), 4) + 0.0
[lam(u) for u in (0.0, 0.25, 0.5, 0.75, 1.0)] # → [20.0, 1.7626, 0.0, -1.7626, -20.0]
The noise guess is the score
Because the loss is denoising score matching, the noise guess estimates the score, the direction in which noisy data gets denser, flipped and scaled:
In words: “the best noise guess points the opposite way to uphill-in-probability, scaled by the noise level.”
With the numbers: if the data were the single value 2.0, noisy values at this level form N(1.6, 0.36), whose score at 1.9 is −(1.9 − 1.6)/0.36 = −0.8333. Times −0.6 gives 0.5, exactly the noise.
In Python:
z, centre, var, sigma = 1.9, 1.6, 0.36, 0.6
# ∇ log N(z; centre, var) = −(z − centre) / var
score = -(z - centre) / var
round(score, 4), round(-sigma * score, 4) # → (-0.8333, 0.5)
For a class-conditional model “the only modification” is that the network also receives the label c: εθ(zλ, c).
Why it matters
The score view is what makes guidance possible: scores of different distributions can be added and scaled, because adding log-probabilities multiplies probabilities. Guidance is built entirely from that move. The paper also notes, and it matters later, that a network's output need not be the gradient of any function.
In code: train_noise_predictor trains the noise guesser; noise_to_score turns a noise guess into a score.
3 Guidance · original
“Unfortunately, straightforward attempts of implementing truncation or low temperature sampling in diffusion models are ineffective.”Ho and Salimans (2022), §3
Everyday picture
Many generators have a “play it safe” knob. A GAN can be fed only typical noise (the truncation trick); a flow model or a language model can sample at low temperature. Each trades variety for samples that look more typical and cleaner.
Tiny example
Lowering the temperature of a bell curve N(0, 1) to 0.5 means raising its density to the power 1/0.5 = 2 and renormalising: e−x²/2 squared is e−x², the bell curve N(0, 0.5). Same centre, spread shrunk from 1 to 0.707: fewer unusual samples.
In Python:
import math
temperature = 0.5
# (e^(−x²/2))^(1/T) = e^(−x²/(2T)): a bell curve with variance T
variance = temperature
round(variance, 4), round(math.sqrt(variance), 4) # → (0.5, 0.7071)
Why it matters
For diffusion models the obvious versions of this knob (scale up the noise guesses, or add less noise while sampling) give blurry, low-quality samples, as the earlier work the paper builds on found. Something else was needed.
3.1 Classifier guidance · original
Everyday picture
At every step, show the half-finished sample to a classifier trained on noisy images, ask how to change it to look more like the label, and nudge it that way as well as along the diffusion model's own direction. The strength of the nudge is w.
Tiny example
In the two-class toy at z = 0.1, a perfect noisy-data classifier says “+” with probability 0.7087, and the slope of its log-probability is +2.59: move right to become more “+”. The class-conditional noise guess for “+” is (0.1 − 1.6)/0.6 = −2.5. With w = 1, the guided guess is −2.5 − 1 × 0.6 × 2.59 = −4.05: a guess of more negative noise, so a step that moves z further right, toward “+”.
In words: “the guided noise guess is the labelled guess minus w times the classifier's push toward the label, rescaled; it is the noise guess you would get for a distribution whose log-probability adds w times the classifier's log-probability.”
With the numbers: −2.5 − 1 × 0.6 × 2.5897 = −4.0538. The classifier's slope comes from Bayes' rule on the two bell curves at centres ±1.6 with variance 0.36: log-odds 8.889 × z, so a slope of 8.889 × (1 − 0.7087) = 2.5897.
In Python:
import math
z, sigma, var = 0.1, 0.6, 0.36
# p(+ | z) by Bayes' rule: a logistic curve in z, with slope k = 2 × 1.6 × 2 / (2 × 0.36)
k = 4 * 1.6 / (2 * var)
p_plus = 1 / (1 + math.exp(-k * z))
round(p_plus, 4), round(k, 3) # → (0.7087, 8.889)
# ∇ log p(+ | z) = k (1 − p)
grad = k * (1 - p_plus)
eps_c = (z - 1.6) / sigma
round(grad, 4), round(eps_c, 4) # → (2.5897, -2.5)
# ε̃ = ε_c − w σ ∇ log p(c | z), with w = 1
round(eps_c - 1 * sigma * grad, 4) # → -4.0538
Using ε̃ in place of the labelled guess samples, approximately, from a sharpened distribution:
In words: “take the class's own distribution and multiply in the classifier's confidence w times: places where the classifier is unsure lose weight, places where it is sure keep theirs.”
With the numbers: compare the ambiguous z = 0.1 with the class centre 1.6, where the classifier is certain. Unguided, z = 0.1 is 0.0439 times as likely as 1.6. With w = 1 that falls to 0.0311 (× 0.7087), and with w = 3 to 0.0156 (× 0.7087³). Ambiguous samples are squeezed out.
In Python:
import math
def density(z, centre, var=0.36):
return math.exp(-(z - centre) ** 2 / (2 * var)) / math.sqrt(2 * math.pi * var)
def p_plus(z):
return density(z, 1.6) / (density(z, 1.6) + density(z, -1.6))
# p(z | c) p(c | z)^w at z = 0.1, relative to the class centre z = 1.6
ratio = density(0.1, 1.6) / density(1.6, 1.6)
[round(ratio * (p_plus(0.1) / p_plus(1.6)) ** w, 4) for w in (0, 1, 3)] # → [0.0439, 0.0311, 0.0156]
Try it: the paper's Figure 2 shows three classes, each a round bell curve in the plane. Drag w and watch each class's guided distribution.
Reading it: each colour is one class (its data a round bell curve of spread 0.7 around a centre 1.0 from the middle, illustrative numbers), and the brightness is the density of the equal mixture of the three guided class distributions p(x | c) p(c | x)w, each renormalised, computed exactly on a grid. At w = 0 it is the plain data distribution: three overlapping blobs that blur into one another in the middle. Drag w up and each class retreats from the others: its mass moves outward, away from where the classifier could confuse it, and bunches into a smaller, lopsided region, no longer a round bell curve. The readout tracks one class. That is the paper's Figure 2, redrawn: higher confidence and lower variety, the Inception-score-up, diversity-down trade in miniature.
Guiding an unconditional model
The paper also notes that guiding an unconditional model with weight w + 1 gives the same thing as guiding a conditional model with weight w, since, by Bayes' rule, p(z | c) is proportional to p(z) p(c | z):
In words: “start from the unlabelled noise guess and push toward the label once more than before: the first push turns the unlabelled guess into the labelled one, the rest is guidance.”
With the numbers: the toy's unlabelled guess at z = 0.1 is −0.9462 (computed in §3.2 below). −0.9462 − 2 × 0.6 × 2.5897 = −4.0538: the same guided guess as before.
In Python:
eps_u, sigma, grad, w = -0.946191, 0.6, 2.589682, 1
# ε_u − (w + 1) σ ∇ log p(c | z)
round(eps_u - (w + 1) * sigma * grad, 4) # → -4.0538
In theory the two are the same; in practice the earlier work got its best results guiding an already conditional model, so the paper stays with that.
Why it matters today
Classifier guidance showed that a diffusion model can be sharpened after training, with one knob. Its costs (a second, noise-trained classifier, and steps along that classifier's gradient) are what the next section removes.
3.2 Classifier-free guidance · original
“Eq. 6 has no classifier gradient present, so taking a step in the ε̃θ direction cannot be interpreted as a gradient-based adversarial attack on an image classifier.”Ho and Salimans (2022), §3.2
Everyday picture
Instead of hiring a critic, teach the artist to draw both ways: sometimes you tell them the subject, sometimes you don't. Then, while drawing, the artist compares “what I'd draw for a dog” with “what I'd draw for anything” and exaggerates the difference. The difference between those two is itself a critic's opinion, recovered without a critic.
Tiny example
At some step the unlabelled guess is 0.2 and the labelled guess is 0.5 (the diffusion lesson's numbers). The label moves the guess by 0.3. With w = 2, go 2 × 0.3 = 0.6 past the labelled guess: 1.1.
Hover or tap a step. Training is on the left, sampling on the right.
Reading it: on the left, training is DDPM's loop with one extra box, the second from the top: with probability puncond the label is replaced by the null label ∅, so the same network learns both the labelled and the unlabelled guess. Everything else (pick a level, noise the picture, score the noise guess, take a gradient step) is unchanged. On the right, sampling starts from pure noise, and at every level calls the network twice, once with the label and once with ∅. The two guesses meet in the mixing box, (1 + w)εc − w εu, which is the whole of classifier-free guidance. The mixed guess predicts a clean picture x̃, and an ordinary sampler step (the paper's own, or DDIM, as its Algorithm 2 notes) moves to the next level. After the last level, x̃ is the output.
The formula
In words: “take the labelled guess, and move w times further in the direction from the unlabelled guess to the labelled one.” Rearranged, it is εu + (1 + w)(εc − εu), which is the diffusion lesson's formula with its weight equal to 1 + w.
With the numbers: (1 + 2) × 0.5 − 2 × 0.2 = 1.5 − 0.4 = 1.1. At w = 0 it is 0.5, the plain labelled guess. The lesson writes the same mix as 0.2 + 3 × (0.5 − 0.2) = 1.1, with its weight 3.
In Python:
eps_c, eps_u = 0.5, 0.2
# the paper: ε̃ = (1 + w) ε_c − w ε_u
[round((1 + w) * eps_c - w * eps_u, 2) for w in (0, 1, 2)] # → [0.5, 0.8, 1.1]
# the lesson: ε_u + w_lesson (ε_c − ε_u), with w_lesson = 1 + w
[round(eps_u + (1 + w) * (eps_c - eps_u), 2) for w in (0, 1, 2)] # → [0.5, 0.8, 1.1]
The implicit classifier
Where does the formula come from? Bayes' rule gives a classifier hidden inside any pair of conditional and unconditional models, pi(c | z) ∝ p(z | c)/p(z). Its slope can be read off the two exact noise guesses:
In words: “the difference between the labelled and the unlabelled noise guess, flipped and rescaled, is exactly the slope of the classifier that Bayes' rule hides in the two models.”
With the numbers: in the two-class toy at z = 0.1, the exact labelled guess is −2.5 and the exact unlabelled guess (averaging the two classes' guesses by their 71% / 29% odds) is −0.9462. Then −(−2.5 + 0.9462)/0.6 = 2.5897: the classifier slope of §3.1, found with no classifier. Plug it into classifier guidance with w = 1 and you get (1 + 1) × (−2.5) − 1 × (−0.9462) = −4.0538, equation 6 exactly.
In Python:
import math
z, sigma, var = 0.1, 0.6, 0.36
centres = [1.6, -1.6]
# the two classes' odds at z, by Bayes' rule
dens = [math.exp(-(z - m) ** 2 / (2 * var)) for m in centres]
odds = [d / sum(dens) for d in dens]
# exact noise guesses: labelled "+", and unlabelled (the odds-weighted average)
eps_c = (z - centres[0]) / sigma
eps_u = sum(p * (z - m) / sigma for p, m in zip(odds, centres))
round(eps_c, 4), round(eps_u, 4) # → (-2.5, -0.9462)
# the implicit classifier's slope, and equation 6 with w = 1
round(-(eps_c - eps_u) / sigma, 4), round(2 * eps_c - eps_u, 4) # → (2.5897, -4.0538)
The paper stresses a subtlety. With exact scores the two methods agree, as above. But a trained network's outputs are not guaranteed to be the slope of any function at all (they need not form a conservative vector field), so the learned difference εθ(z, c) − εθ(z) is in general not the slope of any classifier. Classifier-free guidance is inspired by an implicit classifier, not equal to one, and so it cannot be an adversarial attack on one.
It was also not obvious that it would work: a classifier built from a generative model by Bayes' rule is often worse than one trained directly, and can behave badly when the model is imperfect. The experiments are the evidence.
Try it: the same three classes as Figure 2 above, but now actually sampled: 300 samples of the top class, using the exact labelled and unlabelled noise guesses, equation 6, and 50 DDIM steps (see the DDIM companion). Drag w and compare the dots with the shaded target.
Reading it: the shading is the top class's guided target p(x | c) p(c | x)w from Figure 2, and the dots are what guided sampling actually produces from the same 300 starting noises. At w = 0 the dots roughly match the class's plain bell curve (300 samples and 50 big steps leave small differences in the readout). As w grows the dots move up and away from the other two classes and tighten, just as the shading does, but much further: at w = 4 the samples' centre is about 2.5 from the middle, against 1.45 for the shaded target. The reason is that guidance is applied at every noise level, including the very noisy ones where the classes overlap heavily and the push is strong. So guided sampling only approximates the sharpened distribution of §3.1, and overshoots it toward exaggeration; the paper, for its part, notes that strongly guided ImageNet samples show saturated colours.
Why it matters today
This is the whole method, and it is a few lines of code: drop the label now and then while training; run the network twice and mix while sampling. It needs no classifier, so the “label” can be anything the network can read, such as a sentence.
In code: network_inputs adds the one-hot label with its extra “no label” slot, train_noise_predictor hides labels at random (Algorithm 1), guided_noise is equation 6 in the lesson's form, and ddim_sample applies it at every jump (Algorithm 2 with a DDIM step).
4 Experiments · original
Everyday picture
A proof of concept, not a race: take the architecture and settings the classifier-guidance work had tuned, train one network both ways, and see whether sweeping w traces the same quality-versus-variety curve.
The setup
| Setting | Value |
|---|---|
| Data | class-conditional ImageNet, downsampled to 64 × 64 and 128 × 128 |
| Network | the same architecture and hyperparameters as the earlier classifier-guided models (a U-Net), with continuous-time training |
| Noise levels | λmin = −20, λmax = 20 |
| Training | 400 thousand steps at 64 × 64; 2.7 million at 128 × 128 |
| Sampler variance knob v | 0.3 at 64 × 64; 0.2 at 128 × 128 |
| Sweep | w from 0 to 4 in steps of 0.1 (the tables print 0 to 1 and then 2, 3, 4); FID and Inception score from 50,000 samples each |
Tiny example
The two scores pull in different directions. FID compares the whole cloud of samples with real images, so it punishes lost variety; lower is better. Inception score rewards each sample being confidently one class and the classes being varied; higher is better. Guidance buys confidence with variety, so it should push both numbers up: a better Inception score and a worse FID.
Why it matters
Because the settings were tuned for classifier guidance and one network does double duty (less capacity than a model plus a classifier), any good result here is, if anything, an underestimate.
4.1 Varying the classifier-free guidance strength · original
“…here we clearly see that increasing classifier-free guidance strength has the expected effect of decreasing sample variety and increasing individual sample fidelity.”Ho and Salimans (2022), §4.1
Everyday picture
Turn the dial and the gallery changes character: at w = 0 a varied crowd of dogs, some odd; at w = 3 or 4, every dog is an unmistakable textbook dog, in saturated colours, and they start to look alike. The paper's image figures show exactly this for the malamute class.
Tiny example
| w | FID (lower is better) | Inception score (higher is better) |
|---|---|---|
| 0 (no guidance) | 1.8 | 53.71 |
| 0.1 | 1.55 | 66.11 |
| 1.0 | 12.6 | 170.1 |
| 4.0 | 26.22 | 260.2 |
A little guidance helps both scores (FID 1.8 → 1.55, the best in the table). More guidance keeps raising the Inception score, to 260, while FID climbs to 26: every sample more convincing, the collection less like the real, varied data.
Hover or use the arrow keys to read FID and Inception score along each curve.
Reading it: each line is one trained model, and each point on it is one guidance weight w: the x-axis is Inception score, the y-axis is FID. Guidance grows from left to right along every line, from w = 0 at the far left to w = 4 at the far right, because Inception score rises steadily with w. Look first at the bottom-left corner: the lines dip slightly (FID improves) for the first small step of guidance, then turn up. After that it is a straight trade: every gain in Inception score is paid for in FID. The three lines are three values of puncond (§4.2). This is the paper's Figure 4, redrawn from the numbers in its Table 1.
The paper's best 128 × 128 numbers beat the classifier-guided model it compares with at w = 0.3 (FID 2.43 against 2.97), and at w = 4 beat BigGAN-deep at its best-Inception-score setting on both scores (FID 21.53 against 25, Inception score 421.03 against 253, all from its Table 2).
One sentence in §4.1 reads the wrong way round: it describes “FID monotonically decreasing and IS monotonically increasing with w”, but the paper's own tables and figures show FID increasing with w beyond the first small step, as the chart above does. The trade-off the sentence goes on to describe is the one the numbers show.
Why it matters today
The shape of this curve is why products expose w as a knob and pick a default well above 0: people usually prefer each image to follow the prompt convincingly over the collection matching the data's variety.
In code: guidance_sweep measures the lesson's version of this trade: the share of samples on the requested blob and their spread, at several weights.
4.2 Varying the unconditional training probability · original
Everyday picture
How often should the artist practise without being told the subject? Enough to know what “anything” looks like, but not so often that the labelled drawing suffers.
Tiny example
Out of 400,000 training steps with puncond = 0.1, about 40,000 steps' worth of examples are unlabelled. At 0.5, half of all training goes to the unlabelled task.
In Python:
steps = 400_000
# how many steps' worth of examples train the unlabelled guess
[int(steps * p) for p in (0.1, 0.2, 0.5)] # → [40000, 80000, 200000]
In the chart above, the puncond = 0.5 line sits above the other two all the way along (worse FID at the same Inception score), while 0.1 and 0.2 are almost on top of each other. So a small share of the network's effort, 10 to 20%, is enough for the unlabelled guess that guidance needs. The paper notes the parallel with classifier guidance, where small classifiers were found to suffice.
Why it matters today
Dropping the condition 10 to 20% of the time, replacing it with an empty or null prompt, became the default recipe for training text-conditioned diffusion models.
In code: train_noise_predictor takes puncond as an argument, 0.2 by default.
4.3 Varying the number of sampling steps · original
Everyday picture
More steps give a more careful drawing, but here every step costs two network calls instead of one. A fair race counts calls, not steps.
Tiny example
256 guided steps cost 2 × 256 = 512 network passes. The classifier-guided model the paper compares against used about 256 steps of the same architecture (plus its classifier). So, as the paper says, the fair comparison in network passes is its T = 128 setting, whose best FID (3.02, at w = 0.4) is worse than the classifier-guided 2.97.
In Python:
calls_per_step = 2
# network passes for the paper's three step counts
[T * calls_per_step for T in (128, 256, 1024)] # → [256, 512, 2048]
Hover or use the arrow keys to read FID and Inception score along each curve.
Reading it: as in the chart of §4.1, each line traces w from 0 (left) to 4 (right), with Inception score across and FID up; here the lines are the 128 × 128 model sampled with T = 128, 256 and 1,024 steps. The T = 128 line sits a little above the other two for small w, where FID is lowest, while 256 and 1,024 steps almost coincide. So T = 256 is the good balance the paper names. This is the paper's Figure 5, redrawn from its Table 2. (The printed Figure 5 labels its third line T = 512, but its caption, the text and Table 2 all say 1,024; this redraw uses the table.)
Why it matters today
The doubled cost is real and still paid; a common implementation runs the conditional and unconditional passes together as one batch of two. Cutting the step count, with samplers like DDIM, matters twice as much when every step costs double.
5 Discussion · original
“The most practical advantage of our classifier-free guidance method is its extreme simplicity…”Ho and Salimans (2022), §5
Everyday picture
The paper's own summary of why guidance works: it makes a sample more likely under the labelled model and less likely under the unlabelled one. Aim for “very dog” and away from “generic picture”.
What it says
- Simplicity: one line in training (randomly drop the condition) and one in sampling (mix the two guesses).
- No classifier gradients: the guided sampler's steps need not resemble any classifier's gradient, so good classifier-based scores cannot be blamed on fooling a classifier.
- A negative term: the −w εθ(z) part pushes away from the unlabelled model, which the authors expect to be useful elsewhere.
- Cost: two passes of the full model per step, where a classifier can be small; injecting the condition late in the network might help.
- Diversity: any method that buys quality with diversity may hurt applications where parts of the data are under-represented.
Tiny example: no unconditional model needed, with few classes
If there are only a few classes and their frequencies are known, the unlabelled distribution is just the frequency-weighted sum of the labelled ones, so it need not be trained at all:
In words: “the chance of an example overall is its chance within each class, weighted by how common the class is, added up.”
With the numbers: in the two-class toy at z = 0.1, the classes' densities are 0.0292 and 0.0120; with equal frequencies the overall density is 0.5 × 0.0292 + 0.5 × 0.0120 = 0.0206. The catch: one network pass per class, which is hopeless for ImageNet's 1,000 classes or for free text.
In Python:
import math
def density(z, centre, var=0.36):
return math.exp(-(z - centre) ** 2 / (2 * var)) / math.sqrt(2 * math.pi * var)
z, classes = 0.1, {+1.6: 0.5, -1.6: 0.5}
# p(z | c) for each class, then Σ_c p(z | c) p(c)
[round(density(z, c), 4) for c in classes] # → [0.0292, 0.012]
round(sum(density(z, c) * p for c, p in classes.items()), 4) # → 0.0206
Why it matters today
The “negative term” turned out to be widely useful: image tools let you replace the empty prompt in the unlabelled pass with a negative prompt, steering away from things you do not want. And the diversity concern is a standing one for any strongly guided generator.
6 Conclusion · original
What the paper concludes
Classifier-free guidance increases sample quality while decreasing diversity, like classifier guidance without the classifier. Pure generative diffusion models can therefore maximise classifier-based quality scores without any classifier gradients.
Tiny example
The whole paper fits in the toy: one network that knows both −2.5 (labelled) and −0.9462 (unlabelled) at z = 0.1 can produce the classifier-guided −4.0538 without ever having trained a classifier.
Why it matters today
The authors hoped for “a wider variety of settings and data modalities”. That is what happened: the same two-pass mix steers image, video and audio diffusion models, and flow matching models too, since they also produce a direction from two conditionings that can be mixed.
What changed since 2022
The mixing formula is unchanged in most systems. What surrounds it has moved on:
| In the paper | Common today | Why | Where to read |
|---|---|---|---|
| Class labels | Text prompts, read by cross-attention; the null label is an empty prompt | Ask for anything describable | diffusion lesson |
| Pixel-space ImageNet models | Guidance inside latent diffusion and flow matching models | Cheaper steps, straighter paths | flow matching companion |
| 256 to 1,024 steps, two passes each | DDIM-style samplers with tens of steps, still two passes | Every pass costs time | DDIM companion |
| Unlabelled pass as the anchor | Negative prompts in its place | Steer away from what you do not want | §5 above |
Glossary
Every term with hover guidance on this page, in one place.