Score-Based Generative Modeling through Stochastic Differential Equations, annotated
How to read this page
Nothing on this page assumes you already know the jargon. Three things help:
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation is followed by a table of its symbols, a sentence reading it aloud, the numbers of a tiny example, and the same example in a few lines of Python.
- The pictures are live: press the buttons, drag the sliders, and the readouts say what changed.
Each idea is explained in the same order: an everyday picture, a tiny example you could check by hand, a diagram, then the math, and finally why it still matters.
Which way time runs. In this paper, time t runs from 0 (data) to T = 1 (pure noise), as in the DDPM companion. The flow matching and rectified flow companions run the other way, from noise at 0 to data at 1. Generating here means running time backwards, from 1 down to 0.
The one-pixel toy. A one-pixel image that is either −1 (black) or +1 (white), equally often. The paper's “variance preserving” noising (§3.4, with its settings βmin = 0.1 and βmax = 20) fades each value and mixes in noise. At t = 0.2 the two values have faded to ±0.8114 and each is blurred by noise of variance 0.3416: two overlapping bell curves. Every worked number below uses this toy, and almost all of them sit at the point x = 0.5 at that moment. The paper does the same arithmetic on 3,072-number images; the diffusion lesson builds it from scratch in NumPy on dots in a plane.
Abstract
“Creating noise from data is easy; creating data from noise is generative modeling.”Song et al. (2020), Abstract. Read the original
Everyday picture
Drop ink into a glass of water. Within a minute it has spread into a faint, even cloud: any drop of ink, whatever shape it started as, ends up as the same cloud. Now imagine filming that and playing the film backwards: the cloud gathers itself into a drop. Physics never does this on its own, but it turns out you can play the film backwards if, at every moment, you know one thing: which way is the ink more concentrated? That direction is the score. This paper learns it with a neural network and uses it to run the film backwards, turning noise into images.
What the paper claims
- Adding noise can be written as one continuous process, a stochastic differential equation (SDE), and its reverse is another SDE that needs only the score at each moment.
- The two leading families of the time, score matching with Langevin dynamics and DDPM, are the same idea on two different SDEs.
- New samplers follow: predictor-corrector samplers, and a noise-free probability flow ODE that gives exact likelihoods, a reversible code for every image, and fast adaptive sampling.
- One unconditional model can do class-conditional generation, inpainting and colourisation, with no retraining.
- Record results on CIFAR-10 (FID 2.20, Inception score 9.89, and 2.99 bits per dimension), and the first 1024 × 1024 images from a score-based model.
Why it matters today
This is the paper that gave diffusion models their common language. When later work talks about “the probability flow ODE”, “VP” and “VE” schedules, or solving a diffusion model with an ODE solver, it is using this paper's vocabulary. The flow matching and rectified flow papers both start from its ODE and ask what happens if you choose the road directly.
1 Introduction · original
“We therefore refer to these two model classes together as score-based generative models.”Song et al. (2020), §1
Everyday picture
Two teams were restoring faded photographs. One team (score matching with Langevin dynamics) blurred photos at a ladder of fixed strengths and learned, at each strength, which way “sharper” lies. The other team (DDPM) faded photos in a thousand tiny steps and learned to undo each step. The paper notices that both are snapshots of one continuous fade, and that the continuous version is easier to reason about and to sample from.
Tiny example
Take the toy pixel at +1. Fade it continuously and it drifts toward 0 while noise piles up: at t = 0.2 it is a bell curve centred at 0.8114 with variance 0.3416, and at t = 1 it is centred at 0.0066 with variance 0.99996, indistinguishable from pure noise. The −1 pixel does the mirror image. Reverse the film and the noise has to decide, somewhere along the way, which of the two it will become.
Figure 1, redrawn: the film and its reverse
Reading it: time runs left to right, from data (t = 0) to noise (t = 1); height is the pixel's value. The shading is the crowd pt(x) at each moment: two thin bands at ±1 on the left, which fade inwards and merge into one wide band on the right. Forward SDE: 12 pixels start on the two data values and jitter their way into the noise band, forgetting where they came from. Reverse SDE: 12 pixels start as pure noise on the right and jitter their way leftwards; as they near the left, each settles into one band and ends on −1 or +1 (over many draws, half each; press new random draws). Probability flow ODE: the same starting noise, but no jitter at all. The paths are smooth and never cross, and every one that starts above 0 ends on +1. They barely move for most of the journey and do nearly all their travelling close to t = 0, a slow-then-rushed pace the rectified flow paper later picks on. Given the exact score, both kinds of path trace out the same shading: that is §4.3's claim.
Why it matters
Working in continuous time frees the sampler from the training grid. Any solver, with any number of steps, can follow the reverse process, and the network trained once is reused by all of them.
2 Background · original
The paper first recalls the two methods it unifies, writing both in the same notation so the resemblance jumps out.
2.1 Denoising score matching with Langevin dynamics (SMLD) · original
“For sampling, Song & Ermon 2019 run M steps of Langevin MCMC to get a sample for each pσi(x) sequentially”Song et al. (2020), §2.1
Everyday picture
Picture a hiker in thick fog on a hillside, trying to find the valleys where people live. They can't see the valleys, but they can feel which way the ground slopes. Each step they walk a little downhill toward the crowd, then take a small random stumble. After many steps, hikers who follow this rule end up spread over the valleys in exactly the proportion people live there. That walking rule is Langevin dynamics, and the slope they feel is the score. The trick of SMLD is to start in very thick fog (heavy blur, where the slope is gentle and points toward the middle of everything) and let the fog thin gradually.
Tiny example: learning the slope
A data point x = 1 is blurred with noise of size σ = 0.5 and a noise draw z = 0.4, giving x̃ = 1.2. Around x, the blurred crowd is a bell curve, and its slope at 1.2 points back toward 1 with size (1.2 − 1)/0.5² = 0.8. So the network is told: at 1.2, the answer is −0.8. This is denoising score matching: the target is the score of the blur around one known data point, which you can write down.
In words: “find the network weights that, averaged over every noise level, every data point and every way of blurring it, make the network's slope as close as possible to the slope of the blur around that data point; weight each level by its noise variance so every level counts about equally.”
With the numbers: x̃ = 1 + 0.5 × 0.4 = 1.2; target −(1.2 − 1)/0.25 = −0.8. A guess of −0.5 scores 0.25 × (−0.5 + 0.8)2 = 0.0225. The weight σ2 turns this into (σs + z)2 = (−0.25 + 0.4)2, the same 0.0225: a slope guess is a noise guess in disguise, which is why the diffusion lesson can train on noise and convert with noise_to_score.
In Python:
x, sigma, z = 1.0, 0.5, 0.4
# x̃ = x + σ z: the blurred point
x_tilde = x + sigma * z
x_tilde # → 1.2
# ∇ log p_σ(x̃ | x) = −(x̃ − x)/σ²: the slope of the bell curve around x
target = -(x_tilde - x) / sigma ** 2
round(target, 4) # → -0.8
# σ² ‖s_θ − target‖² for a guess of −0.5
s_theta = -0.5
round(sigma ** 2 * (s_theta - target) ** 2, 4) # → 0.0225
# the same number written as a noise-guessing error, (σ s_θ + z)²
round((sigma * s_theta + z) ** 2, 4) # → 0.0225
Tiny example: one Langevin step
Blur the toy's two values ±1 with σ = 0.6, and stand at x = 0.5. The blurred crowd has two bumps. The score at 0.5 points toward a weighted centre of the two bumps, weighted by how likely each is to have produced 0.5: that centre is tanh(0.5/0.36) = 0.8828 (the +1 bump is far more likely), so the score is (0.8828 − 0.5)/0.36 = 1.0637. One Langevin step with step size ε = 0.05 and a stumble z = 0.3 moves the hiker to 0.6481.
In words: “take a small step along the learned slope, then add a random stumble whose size is the square root of twice the step; repeat M times at this noise level, then move on to a smaller one.”
With the numbers: 0.5 + 0.05 × 1.0637 + √0.1 × 0.3 = 0.5 + 0.0532 + 0.0949 = 0.6481. The step toward the crowd and the stumble are about the same size: the stumble is what keeps a crowd of hikers spread out instead of piling onto one spot.
In Python:
import math
sigma, x = 0.6, 0.5
# the score of half a bell curve at −1 and half at +1, each of spread σ:
# (weighted centre − x) / σ², where tanh gives the weighted centre
s = (-x + math.tanh(x / sigma ** 2)) / sigma ** 2
round(s, 4) # → 1.0637
eps, z = 0.05, 0.3
# x + ε s + √(2ε) z: one Langevin step
round(x + eps * s + math.sqrt(2 * eps) * z, 4) # → 0.6481
Try it: start 2,000 hikers from a wide cloud that leans toward white (a bell curve centred at +1 with spread 1, so 83% start above 0), pick a noise level σ, and take M Langevin steps with the exact score. Watch whether the histogram settles onto the target crowd, which has half its people on each side.
Hover or tap the chart, or focus it and use the arrow keys, to read both curves.
Reading it: the x-axis is the pixel value, the y-axis is how crowded each value is. The solid line is the histogram of 2,000 hikers after M steps; the dashed line is the target, the toy's two values blurred by σ. The step size is ε = 0.1σ², an illustrative choice. Set M to 0 to see the lopsided starting cloud. At σ = 0.5, by about 200 steps the histogram sits on the target, two bumps with half the hikers on each side (the readout measures the gap as the average distance between sorted hikers and sorted exact samples). At σ = 1 the hikers rebalance twice as fast, but the two bumps have blurred into one hump. Now drop σ to 0.2. Each bump sharpens within 20 steps, but 83% of the hikers stay on the white side however many steps you take: in the empty gap between the bumps the slope points back toward whichever bump is nearer, and a stumble of size √(2ε) is far too small to hop the gap. Large noise lets hikers cross between bumps; small noise makes them precise. That is why SMLD starts with a large σ and lowers it level by level, and why §4.2 later uses this very step as its “corrector”.
Why it matters today
Langevin dynamics needs no knowledge of the crowd except its slope, which a network can learn. Every diffusion sampler still has this shape: a step along what the network says, plus, in the stochastic ones, a little fresh noise.
2.2 Denoising diffusion probabilistic models (DDPM) · original
“The objective Eq. 3 described here is Lsimple in Ho et al. 2020, written in a form to expose more similarity to Eq. 1.”Song et al. (2020), §2.2
Everyday picture
DDPM fades a photo in a thousand small steps, each one shrinking it slightly toward grey and sprinkling in a little noise, and learns to undo one step at a time. The paper rewrites DDPM's loss so that it looks exactly like SMLD's: the only differences are how the blur is made (shrink and add, rather than just add) and the weight of each level.
Tiny example
This paper's αi is DDPM's running product ᾱ (the signal share left after i steps). With αi = 0.64, the data point x = 1 and noise z = 0.5 become x̃ = 0.8 × 1 + 0.6 × 0.5 = 1.1. The slope of the blur around the shrunken point 0.8 is −(1.1 − 0.8)/0.36 = −0.8333, which is the noise flipped and divided by its size: −0.5/0.6.
In words: “the same recipe as SMLD, but the blur at step i shrinks the data by √αi before adding noise of variance 1 − αi, and each step is weighted by that noise variance.”
With the numbers: x̃ = 1.1, target −0.8333, weight 0.36. The paper points out that both weights, σi2 and 1 − αi, are proportional to one over the target's typical squared size (here that size is 1/0.36 in one dimension), so both losses give every noise level about an equal say.
In Python:
import math
alpha_i, x, z = 0.64, 1.0, 0.5
# x̃ = √α_i x + √(1 − α_i) z: shrink, then add noise
x_tilde = math.sqrt(alpha_i) * x + math.sqrt(1 - alpha_i) * z
round(x_tilde, 4) # → 1.1
# ∇ log p_α(x̃ | x) = −(x̃ − √α_i x)/(1 − α_i)
target = -(x_tilde - math.sqrt(alpha_i) * x) / (1 - alpha_i)
round(target, 4) # → -0.8333
# the same as −z/√(1 − α_i): the noise, flipped and rescaled
round(-z / math.sqrt(1 - alpha_i), 4) # → -0.8333
DDPM samples by walking its chain backwards, one step at a time, with a fresh wobble each step; the paper calls this ancestral sampling:
In words: “nudge the point along the score by this step's noise dose, undo this step's shrink, and add a small fresh wobble.”
With the numbers: with β = 0.02, x = 1.1, score −0.8333 and wobble z = 0.1: (1.1 − 0.0167)/0.98995 + 0.1414 × 0.1 = 1.0943 + 0.0141 = 1.1085.
In Python:
import math
beta_i, x_i, s, z = 0.02, 1.1, -0.8333, 0.1
# (x_i + β_i s) / √(1 − β_i) + √β_i z
x_prev = (x_i + beta_i * s) / math.sqrt(1 - beta_i) + math.sqrt(beta_i) * z
round(x_prev, 4) # → 1.1085
Why it matters today
Written this way, “guess the noise” (DDPM) and “guess the slope” (SMLD) are the same training problem up to a rescaling. The diffusion lesson makes the same point: its train_noise_predictor learns the noise, its noise_to_score turns that into the score, and its ddpm_step is the ancestral step above. The DDPM companion walks through the original paper.
3 Score-based generative modeling with SDEs · original
“Perturbing data with multiple noise scales is key to the success of previous methods.”Song et al. (2020), §3
Everyday picture
SMLD used a ladder of fog levels, DDPM a thousand small steps. The paper asks: why not infinitely many, a fog that thickens smoothly with time? Then “the fog level” becomes a clock, and everything else (training, sampling, likelihood) is stated at every tick of that clock at once.
Tiny example
DDPM with 1,000 steps of βi each is a staircase; the continuous version is the ramp the staircase approximates. With the paper's DDPM schedule, the staircase's centre factor √αi and the ramp's e−½∫β never differ by more than 0.001 over all 1,000 steps (the paper plots the same check in its Figure 5).
Figure 2, redrawn: the whole framework
Hover or tap a block. Start with data on the left.
Reading it: the top lane runs left to right: a fixed forward SDE, with no learning in it, fades data into noise. The two lanes below run right to left, turning noise back into data. Both need one ingredient, the score at every moment, which comes from the score network in the middle; the network is trained (eq. 7) on data faded by the forward SDE, which is why a wire runs from the data to “train”. The reverse SDE (§3.2) keeps adding fresh noise as it goes; the probability flow ODE (§4.3) adds none. Given the same score, both produce the same crowd at every moment.
Why it matters
The picture separates three choices that earlier methods had tied together: how to add noise (which SDE), how to learn the score (which loss), and how to sample (which solver). The rest of the paper explores each one.
3.1 Perturbing data with SDEs · original
“This diffusion process can be modeled as the solution to an Itô SDE”Song et al. (2020), §3.1
Everyday picture
A leaf floating down a river moves for two reasons: the current carries it steadily (the drift), and eddies knock it about at random (the diffusion). A stochastic differential equation is a rule that says, at every moment, how strong each is. Over a short time dt the current moves the leaf by (current) × dt, and the eddies move it by a random amount whose spread grows like √dt, the signature of a Wiener process.
Tiny example
The toy pixel sits at x = 0.5 at t = 0.2, where the paper's schedule gives β = 4.08. Over dt = 0.01 the current pulls it toward 0 by ½ × 4.08 × 0.5 × 0.01 = 0.0102, and an eddy with z = 1 pushes it by √4.08 × √0.01 = 0.2020. It ends at 0.6918. Over one short step the random kick is twenty times the steady pull: noise dominates locally, and the pull only wins over many steps.
In words: “over a tiny time dt, the point moves by the drift times dt, plus the noise strength times a tiny random kick.” Simulating it with small steps like this is an Euler-Maruyama step.
With the numbers: the variance preserving choice (§3.4) is f = −½β(t)x and g = √β(t). At t = 0.2, x = 0.5: f = −1.02, g = 2.0199, and 0.5 − 1.02 × 0.01 + 2.0199 × 0.1 × 1 = 0.6918.
In Python:
import math
t, x, dt, z = 0.2, 0.5, 0.01, 1.0
# β(t) = β_min + t (β_max − β_min): the paper's schedule
beta = 0.1 + t * (20 - 0.1)
round(beta, 2) # → 4.08
# the VP choice: f(x, t) = −½ β(t) x and g(t) = √β(t)
f, g = -0.5 * beta * x, math.sqrt(beta)
round(f, 4), round(g, 4) # → (-1.02, 2.0199)
# dw over a step dt is √dt z
round(x + f * dt + g * math.sqrt(dt) * z, 4) # → 0.6918
Why it matters today
Nothing in the forward SDE is learned. It is a fixed recipe for destroying information, chosen so that its end point, pT, is plain noise you know how to draw. The lesson's add_noise jumps straight to any time along it.
3.2 Generating samples by reversing the SDE · original
“A remarkable result from Anderson 1982 states that the reverse of a diffusion process is also a diffusion process, running backwards in time and given by the reverse-time SDE”Song et al. (2020), §3.2
Everyday picture
Running the river backwards is not just flipping the current. The eddies still scatter the leaf; to end up back where leaves really started, you must add a second current that pushes toward where leaves are more crowded, strong enough to beat the scattering. That extra current is the score, scaled by the eddies' strength squared.
Tiny example
At t = 0.2, x = 0.5 in the toy, the score is (0.6733 − 0.5)/0.3416 = 0.5073: the two faded bumps sit at ±0.8114, and at 0.5 the weighted centre is 0.6733. The reverse drift is −1.02 − 4.08 × 0.5073 = −3.0896. Time runs backwards, so dt = −0.01, and the point moves up by 0.0309, toward the +1 bump.
In words: “going backwards in time, move by the forward drift minus the noise strength squared times the score, and keep adding random kicks of the same strength as before.”
With the numbers: m = e−0.209 = 0.8114, v = 1 − e−0.418 = 0.3416, score 0.5073, drift −3.0896. One step with dt = −0.01 and the random kick set to 0: 0.5 + 0.0309 = 0.5309.
In Python:
import math
t, x = 0.2, 0.5
beta = 0.1 + t * (20 - 0.1)
# ∫₀ᵗ β(s) ds, then each value's faded centre m and noise variance v
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
m, v = math.exp(-B / 2), 1 - math.exp(-B)
round(m, 4), round(v, 4) # → (0.8114, 0.3416)
# the score of the toy's p_t: (weighted centre − x) / v
score = (-x + m * math.tanh(m * x / v)) / v
round(score, 4) # → 0.5073
# f − g² ∇log p_t: the reverse drift
drift = -0.5 * beta * x - beta * score
round(drift, 4) # → -3.0896
# one step back in time (dt = −0.01), with the random kick set to 0
round(x + drift * -0.01, 4) # → 0.5309
Press reverse SDE in Figure 1 to watch this rule, applied 400 times with fresh kicks, turn noise into ±1.
Why it matters today
The only unknown in the reverse SDE is the score. Everything else, f and g, was chosen by hand. So generative modelling reduces to one supervised learning problem: estimate ∇x log pt(x) for every t.
3.3 Estimating scores for the SDE · original
“Note that Eq. 7 uses denoising score matching, but other score matching objectives, such as sliced score matching … and finite-difference score matching … are also applicable here.”Song et al. (2020), §3.3
Everyday picture
Training is the SMLD and DDPM recipe with a continuous dial: pick a random moment t, fade a real example to that moment in one jump, and grade the network's slope against the slope of the fade around that one example. The network sees the moment t as an input, so one network covers every moment.
Tiny example
The white pixel x(0) = 1 faded to t = 0.2 with noise z = 0.5 lands at 0.8114 + 0.5845 × 0.5 = 1.1036. The target slope is −z/√v = −0.8554. A guess of −0.7, weighted by λ = v = 0.3416, scores 0.0083.
In words: “draw a moment uniformly, a real example, and a faded version of it at that moment; grade the network's slope against the slope of the fade around that example; weight moments so they count equally; average.”
With the numbers: x(t) = 1.1036, target −(1.1036 − 0.8114)/0.3416 = −0.8554, loss 0.3416 × (−0.7 + 0.8554)2 = 0.0083. With enough data and a big enough network, the paper notes, the best possible answer is the true score of the whole crowd, not of the one example: averaging over every example that could have faded to 1.1036 takes care of that.
In Python:
import math
t, x0, z = 0.2, 1.0, 0.5
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
m, v = math.exp(-B / 2), 1 - math.exp(-B)
# x(t) drawn from the fade around x(0): N(m x(0), v)
xt = m * x0 + math.sqrt(v) * z
round(xt, 4) # → 1.1036
# ∇ log p_0t(x(t) | x(0)) = −(x(t) − m x(0)) / v
target = -(xt - m * x0) / v
round(target, 4) # → -0.8554
# λ(t) = v, the usual weighting, times the squared error of a guess of −0.7
round(v * (-0.7 - target) ** 2, 4) # → 0.0083
Why it matters today
To use this loss you must be able to jump straight to any moment, which needs the fade's formula (its perturbation kernel). For the SDEs of §3.4 that formula is a simple bell curve. That is why the paper sticks to drifts that simply rescale x, and why the flow matching and rectified flow papers after it choose paths whose fades are bell curves too: it keeps training to one network call per example.
3.4 Examples: VE, VP and sub-VP SDEs · original
“Due to this difference, we hereafter refer to Eq. 9 as the Variance Exploding (VE) SDE, and Eq. 11 the Variance Preserving (VP) SDE.”Song et al. (2020), §3.4
Everyday picture
Two ways to hide a photo in static. Variance exploding: never touch the photo, just pile on louder and louder static until the photo is a whisper in a storm. Variance preserving: turn the photo's volume down as you turn the static up, so the total loudness stays the same. The paper shows SMLD is the first, DDPM the second, and invents a third, sub-VP, which adds a little less static than VP at every moment.
Tiny example
The toy's pixels, ±1 equally often, have variance exactly 1. Under VE (with σmin = 0.01 and σmax = 2, the largest distance between the two data values, following the rule the paper borrows from its predecessor) the variance climbs to 1 + 4 = 5. Under VP it stays at 1 the whole way. Under sub-VP it dips, to 0.7751 at t = 0.2, and comes back to 1 at the end.
In words: “no drift at all; the noise strength is whatever makes the total noise variance grow along σ2(t).” SMLD's geometric noise levels become σ(t) = σmin(σmax/σmin)t (Appendix C).
With the numbers: at t = 0.5, σ = 0.01 × 2000.5 = 0.1414, and d[σ2]/dt = 2σ2 ln 200, so the noise strength is 0.4604.
In Python:
import math
s_min, s_max, t = 0.01, 2.0, 0.5
# σ(t) = σ_min (σ_max/σ_min)^t: SMLD's noise levels, made continuous
sigma = s_min * (s_max / s_min) ** t
round(sigma, 4) # → 0.1414
# d[σ²]/dt = 2 σ² ln(σ_max/σ_min); the noise strength is its square root
g = math.sqrt(2 * sigma ** 2 * math.log(s_max / s_min))
round(g, 4) # → 0.4604
In words: “shrink the point toward zero at rate ½β(t) while adding noise of strength √β(t): the shrinking removes exactly as much variance as the noise adds, when the variance is 1.” DDPM's discrete chain, with 1,000 steps, becomes this with β(t) = 0.1 + 19.9t.
With the numbers: the fade from x(0) to t = 0.2 is a bell curve with centre e−0.209x(0) = 0.8114 x(0) and variance 1 − e−0.418 = 0.3416 (Appendix B, equation 29). A crowd of variance 1 keeps variance 1; a crowd of variance 4 is pulled down to 2.9751 by t = 0.2, on its way to 1.
In Python:
import math
t = 0.2
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
# the VP fade from x(0): centre e^(−B/2) x(0), variance 1 − e^(−B)
round(math.exp(-B / 2), 4), round(1 - math.exp(-B), 4) # → (0.8114, 0.3416)
# Σ_VP(t) = 1 + e^(−B)(Σ(0) − 1), for data of variance 1 and of variance 4
for sigma0 in (1.0, 4.0):
print(round(1 + math.exp(-B) * (sigma0 - 1), 4)) # → 1.0 2.9751
In words: “the same shrinking as VP, but the noise is turned down by a factor that is 0 at the start and approaches 1 later, so early on much less noise is added.”
With the numbers: at t = 0.2 the sub-VP noise strength is √(4.08 × (1 − e−0.836)) = 1.5204, against VP's 2.0199.
In Python:
import math
t = 0.2
beta = 0.1 + t * (20 - 0.1)
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
# sub-VP noise strength √(β(1 − e^(−2B))) against VP's √β
round(math.sqrt(beta * (1 - math.exp(-2 * B))), 4), round(math.sqrt(beta), 4) # → (1.5204, 2.0199)
Appendix B works out how the variance of the whole crowd evolves under each. For VP and sub-VP:
In words: “VP pulls any starting variance toward 1 exponentially fast; sub-VP does the same but always sits at or below VP, and both reach 1 in the end.”
With the numbers: for the toy (Σ(0) = 1) at t = 0.2: VP gives 1 + e−0.418 × 0 = 1; sub-VP gives 1 + 0.4334 − 0.6584 = 0.7751.
In Python:
import math
t, sigma0 = 0.2, 1.0
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
vp = 1 + math.exp(-B) * (sigma0 - 1)
sub_vp = 1 + math.exp(-2 * B) + math.exp(-B) * (sigma0 - 2)
round(vp, 4), round(sub_vp, 4) # → (1.0, 0.7751)
Hover or tap the chart, or focus it and use the arrow keys, to read the three curves.
Reading it: the x-axis is time from data (0) to noise (1); the y-axis is the variance of the whole crowd at that moment. The three lines are VE (σmin = 0.01, σmax = 2), VP and sub-VP, all with the paper's β schedule, computed from Appendix B's formulas. At Σ(0) = 1 (the toy) the VP line is perfectly flat: that is the name. Slide the data variance to 3: VP now falls, and to 0.25: it rises, but either way it ends at 1. Sub-VP always stays under VP and meets it at the end. VE ignores the data's variance and keeps climbing, “exploding”, off the top of the chart: it ends at Σ(0) + 4, five times the toy's variance.
Why it matters today
These names stuck. “VP” is the DDPM family, including the lesson's linear_schedule; “VE” is the SMLD family; and Table 2 below shows sub-VP gave the best likelihoods. Every one of them is a bell-curve fade, so all train with the one-jump loss of §3.3.
4 Solving the reverse SDE · original
Everyday picture
Once the score network is trained, sampling is a driving problem: the reverse SDE is the road, and a numerical solver is the driver. This section tries four drivers: a general-purpose one, a new one that copies the forward recipe, a driver who stops to check the map after every step (predictor-corrector), and a driver on a noise-free road (the probability flow ODE).
4.1 General-purpose numerical SDE solvers · original
“Ancestral sampling, the sampling method of DDPM (Eq. 4), actually corresponds to one special discretization of the reverse-time VP SDE (Eq. 11)”Song et al. (2020), §4.1
Everyday picture
Any recipe for stepping through an SDE will do, the simplest being the Euler-Maruyama step of §3.1 run backwards. The paper adds a “reverse diffusion” sampler: take the forward SDE's own step formula and mirror it. It is easy to write down for any SDE, where deriving DDPM-style ancestral steps can be hard.
Tiny example
With the §2.2 numbers (β = 0.02, x = 1.1, score −0.8333, wobble 0.1), DDPM's ancestral step lands at 1.10847 and the reverse diffusion step at 1.10853. Two different discretisations of the same SDE, agreeing to four decimals for a small step (Appendix E shows they agree as β → 0).
In words: “undo the forward step's shrink to first order, step along the score by the noise dose, then add a fresh kick of the forward step's own size.” This is the predictor of the paper's Algorithm 3.
With the numbers: (2 − 0.98995) × 1.1 − 0.02 × 0.8333 + 0.1414 × 0.1 = 1.11105 − 0.01667 + 0.01414 = 1.10853, against ancestral sampling's 1.10847.
In Python:
import math
beta, x, s, z = 0.02, 1.1, -0.8333, 0.1
# DDPM's ancestral step, equation 4
ancestral = (x + beta * s) / math.sqrt(1 - beta) + math.sqrt(beta) * z
# the reverse diffusion step: Algorithm 3's predictor
reverse_diffusion = (2 - math.sqrt(1 - beta)) * x + beta * s + math.sqrt(beta) * z
round(ancestral, 5), round(reverse_diffusion, 5) # → (1.10847, 1.10853)
Why it matters today
The paper's Table 1 shows reverse diffusion slightly beating ancestral sampling in every column (4.79 against 4.98 for the VE model with 1,000 steps, for example). The deeper point is the one in the quote: DDPM's sampler was an SDE solver all along, so better solvers could replace it without retraining. That is the door DDIM and every later fast sampler walked through.
4.2 Predictor-corrector samplers · original
“Specifically, at each time step, the numerical SDE solver first gives an estimate of the sample at the next time step, playing the role of a “predictor”.”Song et al. (2020), §4.2
Everyday picture
A hiker descending in stages: at each stage, first stride down to the next contour line (predictor), then shuffle around a little on that contour until the group is spread out the way people on that contour really are (corrector). The corrector is §2.1's Langevin step, which needs only the score, and the score is exactly what we have. This kind of alternation is borrowed from numerical methods called predictor-corrector methods.
Tiny example
At t = 0.2 and x = 0.5, the score is 0.5073. With the paper's “signal-to-noise ratio” r = 0.16 (the value it uses for VE models), the corrector's step size comes out as ε = 0.131, and the pull alone moves the point to 0.5665. The ratio of that pull to the size of the random kick is r times √αi: 0.1298.
Hover or tap a block. Start with start at the top.
Reading it: read from the top. The sampler starts from noise at the last level and loops: one predictor step moves it down one level (any SDE solver: ancestral, reverse diffusion, or the probability flow of §4.3), then M corrector steps shuffle it at that level. The “no” arrow returns to the predictor until level 0. The paper's two older methods are special cases: SMLD is “corrector only” (the predictor does nothing but switch the level), and DDPM is “predictor only” (M = 0). The paper uses M = 1 almost everywhere.
The corrector's step size is set so that the pull and the random kick keep a fixed ratio r (Appendix G, Algorithm 5):
In words: “choose the Langevin step so that the score's pull is r times the size of the random kick (scaled by the signal share √αi for VP models).” A bigger score means a smaller step, which keeps the corrector from overshooting where the crowd is steep.
With the numbers: αi = 0.81142 = 0.6584, ‖z‖ replaced by √d = 1 in one dimension (as Appendix G suggests for small batches), ‖g‖ = 0.5073: ε = 2 × 0.6584 × (0.16/0.5073)2 = 0.131. The pull is 0.131 × 0.5073 = 0.0665 and the kick √(2 × 0.131) = 0.5119; their ratio is 0.1298 = 0.16 × 0.8114.
In Python:
import math
t, x, r = 0.2, 0.5, 0.16
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
m, v = math.exp(-B / 2), 1 - math.exp(-B)
# α_i: DDPM's signal share of variance at this level
alpha_i = m * m
# g: the score at the current point
g = (-x + m * math.tanh(m * x / v)) / v
# ε = 2 α_i (r ‖z‖ / ‖g‖)², with ‖z‖ = √d = 1 in one dimension
eps = 2 * alpha_i * (r * 1.0 / abs(g)) ** 2
round(eps, 4) # → 0.131
# the pull alone (the kick set to 0)
round(x + eps * g, 4) # → 0.5665
# the pull's size over the kick's size √(2ε) ‖z‖
round(eps * abs(g) / math.sqrt(2 * eps), 4) # → 0.1298
The results
| Predictor | P1000 | P2000 | C2000 | PC1000 |
|---|---|---|---|---|
| reverse diffusion | 4.79 | 4.74 | 20.43 | 3.60 |
| probability flow | 15.41 | 10.54 | 20.43 | 3.51 |
Compare the last three columns: they cost the same number of network calls (2,000). Doubling the predictor steps (P2000) helps a little; spending everything on the corrector (C2000) is much worse; one of each (PC1000) is best, for every predictor. The corrector-only number belongs to no predictor, so the paper prints it once across the rows. For the variance preserving (DDPM) model the gains are smaller: reverse diffusion goes from 3.21 to 3.18.
Why it matters today
The lasting idea is the split itself: some network calls move the samples down the levels, others fix the crowd at a level, and the two can be traded against each other for the same budget. The next section opens the other route, dropping the noise altogether.
4.3 Probability flow and connection to neural ODEs · original
“For all diffusion processes, there exists a corresponding deterministic process whose trajectories share the same marginal probability densities {pt(x)} as the SDE.”Song et al. (2020), §4.3
Everyday picture
Think of a crowd leaving a stadium. Each person wanders a bit at random, but the crowd as a whole flows smoothly out of the gates. You could replace every wanderer with a sleepwalker who follows the crowd's average flow exactly: individuals take different paths, but a photograph of the crowd at any moment looks the same. The probability flow ODE is that crowd of sleepwalkers. It is an ordinary differential equation: no randomness at all after the starting point.
Tiny example
At t = 0.2, x = 0.5, the reverse SDE's drift was −3.0896 plus random kicks. The probability flow uses only half the score term and no kicks: −1.02 − ½ × 4.08 × 0.5073 = −2.0548. Why half? The kicks spread the crowd out; the SDE needs the extra half of the score pull to undo that spreading. Without kicks, there is nothing to undo.
In words: “move by the forward drift minus half the noise strength squared times the score, and nothing else.” It runs forwards (data to noise) or backwards (noise to data) equally well.
With the numbers: drift −2.0548; one step back with dt = −0.01 moves 0.5 to 0.5205, a smaller move than the SDE's 0.5309 (before its kick).
In Python:
import math
t, x = 0.2, 0.5
beta = 0.1 + t * (20 - 0.1)
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
m, v = math.exp(-B / 2), 1 - math.exp(-B)
score = (-x + m * math.tanh(m * x / v)) / v
# f − ½ g² ∇log p_t: half the reverse SDE's score pull, and no kicks
drift = -0.5 * beta * x - 0.5 * beta * score
round(drift, 4) # → -2.0548
round(x + drift * -0.01, 4) # → 0.5205
Press probability flow ODE in Figure 1: the paths are smooth, never cross, and fill the same shading as the SDE's jittery ones. Because the ODE is smooth, a solver can take much bigger steps along it:
Hover or tap the chart, or focus it and use the arrow keys, to compare the two samplers at each step count.
Reading it: the x-axis is the number of equal time steps (network calls per sample), on a log scale; the y-axis is how far 1,500 samples are from the toy's data, measured as the average distance between the sorted samples and a sorted list of half −1s and half +1s. Lower is better, the floor near 0.01 is the tiny noise left at the paper's stopping time ε = 0.001, and anything worse than 1.2 is drawn at the top edge: with 2 or 3 big steps the SDE's kicks throw samples far off. The solid line is the reverse SDE (Euler-Maruyama steps); the dashed line is the probability flow ODE (Euler steps), both with the exact score. With 10 steps the ODE is already within 0.1 of the data while the SDE is still about 0.5 away; the SDE needs roughly ten times as many steps to catch up. This is an illustrative toy with a perfect score: with a learned score and 1,000 steps, the paper found the SDE samplers better (Table 1 above, by a wide margin for the VE model), and it is at small step counts that the ODE shines.
Exact likelihood
Because the ODE is reversible and smooth, it can also say exactly how probable any image is, as other neural ODEs do. Follow the image forward to noise, and add up how much the flow stretched space on the way: the divergence of the drift, which measures how fast nearby points spread apart. Tiny example: data drawn from a bell curve of variance 0.25, and the point x(0) = 0.3. The flow carries it to 0.6 (the crowd's width doubles), space is stretched by a factor of 2 in total (ln 2 = 0.6931 added up), and the formula returns log p0(0.3) = −0.4058, exactly the bell curve's own value.
In words: “the log-probability of a data point is the log-probability of the noise it flows to, plus the total amount the flow stretched space along the way.” (Appendix D.2, equation 39; stretching makes the noise end less dense, so it has to be added back.)
With the numbers: x(T) = 0.6, log pT(0.6) = −1.0989, total stretch 0.6931, sum −0.4058; the data's own bell curve gives −0.4058 at 0.3.
In Python:
import math
# data: a bell curve of variance 0.25, faded by the VP SDE as before
s2 = 0.25
def B(t):
return 0.1 * t + 0.5 * t * t * (20 - 0.1)
def beta(t):
return 0.1 + t * (20 - 0.1)
def Sigma(t):
# the crowd's variance at time t (Appendix B)
return 1 + math.exp(-B(t)) * (s2 - 1)
def log_normal(x, var):
return -0.5 * math.log(2 * math.pi * var) - x * x / (2 * var)
# follow the probability flow from x(0) = 0.3 to t = 1, adding up ∇·f̃ on the way
x, total, n = 0.3, 0.0, 10_000
h = 1 / n
for i in range(n):
# for a bell-curve crowd, f̃ = k x with k = −½β(1 − 1/Σ), so ∇·f̃ = k
k = -0.5 * beta(i * h) * (1 - 1 / Sigma(i * h))
total += k * h
x += k * x * h
round(x, 4), round(total, 4) # → (0.6, 0.6931)
round(log_normal(x, Sigma(1.0)), 4) # → -1.0989
# log p_0(x(0)) = log p_T(x(T)) + ∫ ∇·f̃ dt
round(log_normal(x, Sigma(1.0)) + total, 4) # → -0.4058
# the exact answer, from the data's own bell curve
round(log_normal(0.3, s2), 4) # → -0.4058
| Model | NLL | FID (ODE) |
|---|---|---|
| Flow++ (a normalizing flow) | 3.29 | - |
| DDPM (Lsimple), as published | ≤ 3.75* | 3.17 |
| the same DDPM, likelihood via the ODE | 3.28 | 3.37 |
| DDPM, continuous training (VP) | 3.21 | 3.69 |
| DDPM, continuous training (sub-VP) | 3.05 | 3.56 |
| DDPM++ continuous (deep, sub-VP) | 2.99 | 2.92 |
Read down: the very same trained DDPM scores 3.28 instead of its bound of 3.75 once the likelihood is computed exactly; training with the continuous loss helps; sub-VP beats VP every time; and the best model reaches 2.99, a record for this kind of evaluation, without ever being trained to maximise likelihood. The paper also notes three other gifts of the ODE: every image has a latent code (where it flows to at t = T) that can be interpolated or rescaled; that code is uniquely identifiable, since the forward SDE has nothing learned in it, so two differently built models trained well give nearly the same code (Appendix D.5 checks this); and a black-box adaptive ODE solver can cut network calls by over 90% without visibly harming samples (Figure 3 of the paper).
Why it matters today
This section is the root of fast diffusion sampling. DDIM, derived independently, turns out to step along this same ODE (see the DDIM companion, Proposition 1, and the lesson's ddim_step). And once sampling is “solve an ODE”, the next question is obvious: why not train the ODE's road directly, and make it straight? That is the flow matching and rectified flow papers, whose samplers are the lesson's euler_step.
4.4 Architecture improvements · original
“Surprisingly, we can achieve better FID than the previous best conditional generative model without requiring labeled data.”Song et al. (2020), §4.4
Everyday picture
A better recipe also deserves a better oven. The authors rebuilt DDPM's U-Net with parts borrowed from the best image GANs, searched over the combinations, and named the winners after the method each served: NCSN++ for the VE SDE and DDPM++ for VP and sub-VP.
What changed in the network
- Smoother up- and down-sampling of images (anti-aliased filters, as in StyleGAN2).
- Skip connections scaled down by 1/√2, which keeps the sum of two paths at the same size as each.
- BigGAN's residual blocks, and 4 of them per resolution instead of 2 (8 in the “deep” models).
- For continuous training, the time t enters through random Fourier features instead of DDPM's sinusoidal step embedding.
| Model | FID | IS |
|---|---|---|
| StyleGAN2-ADA, conditional (given the class) | 2.42 | 10.14 |
| StyleGAN2-ADA, unconditional | 2.92 | 9.83 |
| DDPM | 3.17 | 9.46 |
| DDPM++ continuous (deep, VP) | 2.41 | 9.68 |
| NCSN++ continuous (deep, VE) | 2.20 | 9.89 |
On FID the best VE model beat every listed model, including a GAN that was told which class to draw. The trade-off the paper reports is consistent: VE gave the best samples, VP and sub-VP the best likelihoods, so “practitioners likely need to experiment with different SDEs”. The same recipe produced the first 1024 × 1024 face images from a score-based model (Appendix H.3), with visible flaws such as imperfect facial symmetry.
Why it matters today
Much of the jump from DDPM's 3.17 to 2.20 is architecture and training, not the SDE view itself; the paper is careful to separate the two (Table 3 lists the old-objective NCSN++ at 2.45). The lesson that bigger, better-conditioned denoisers matter carried straight on into transformer denoisers.
5 Controllable generation · original
“The continuous structure of our framework allows us to not only produce data samples from p0, but also from p0(x(0) | y) if pt(y | x(t)) is known.”Song et al. (2020), §5
Everyday picture
You want not just any photo but a photo of a horse. The score says “toward any photo”; a classifier that can recognise horses in noisy images says “toward more horse”. Add the two arrows and follow the sum. Because the score is a slope of a logarithm, and the log of a product is a sum, Bayes' rule becomes plain addition of arrows. This is the idea behind classifier guidance.
Tiny example
Give the toy a label: y = +1 means “white”. At t = 0.2 and x = 0.5, a perfect noisy-image classifier says the pixel is white with probability 0.9149; the slope of its log is 0.4042. The unconditional score was 0.5073. Their sum, 0.9115, is exactly the score of the white bump alone, (0.8114 − 0.5)/0.3416: the classifier has deleted the black bump from the map.
In words: “the reverse SDE of §3.2, with the score replaced by the score plus the slope of the classifier's log-probability for the wanted label.” The replacement is exact because of Appendix I's one-line identity: ∇ log pt(x | y) = ∇ log pt(x) + ∇ log pt(y | x).
With the numbers: pt(white | 0.5) = 0.9149, slope 0.4042, sum 0.9115 = (0.8114 − 0.5)/0.3416. The reverse drift becomes −1.02 − 4.08 × 0.9115 = −4.7388, a stronger pull toward white than the unconditional −3.0896.
In Python:
import math
t, x = 0.2, 0.5
beta = 0.1 + t * (20 - 0.1)
B = 0.1 * t + 0.5 * t * t * (20 - 0.1)
m, v = math.exp(-B / 2), 1 - math.exp(-B)
# ∇log p_t(x): the unconditional score
uncond = (-x + m * math.tanh(m * x / v)) / v
# p_t(y = white | x): the share of the crowd at x that came from +1
p_white = 1 / (1 + math.exp(-2 * m * x / v))
round(p_white, 4) # → 0.9149
# ∇log p_t(y = white | x): the slope of the classifier's log-probability
grad = 2 * m / v * (1 - p_white)
round(uncond, 4), round(grad, 4) # → (0.5073, 0.4042)
# their sum is the score of the white bump alone, (m − x)/v
round(uncond + grad, 4), round((m - x) / v, 4) # → (0.9115, 0.9115)
# the conditional reverse drift f − g² (score + slope)
round(-0.5 * beta * x - beta * (uncond + grad), 4) # → -4.7388
Hover or tap the chart, or focus it and use the arrow keys, to read the three arrows at any pixel value.
Reading it: the x-axis is the pixel value; the y-axis is the size of an arrow (positive means “push up, toward white”). At t = 0.15 the crowd still has two bumps, centred at ±0.89. The solid curve is the unconditional score: between the bumps it points toward whichever is nearer (positive just above 0, negative just below), and beyond them it points back in. The dashed curve is the classifier's slope for “white”: always positive, largest where the pixel looks black, and fading to 0 where the classifier is already sure it is white. The dotted line is their sum, the conditional score: a straight line that crosses 0 exactly at the white bump's centre, so every pixel is pushed toward white. Drag t past about 0.25 and the two bumps merge into one hump: the plain score now points toward 0 everywhere, while the conditional score still points at white. Toward t = 1 everything flattens, because at pure noise there is little left to push toward.
Imputation and colourisation
Inpainting is conditioning on part of the image itself. Tiny example: a two-pixel image where the first pixel is known. At each reverse step, fade the known pixel to the current noise level with the forward SDE (you know its fade exactly), paste it in, and let the reverse SDE update only the unknown pixel using the score of the pasted-together image (Appendix I.2). For colourisation, the known part is the grey level, which is a mix of all three colour channels; an orthogonal change of colour axes (one axis is the grey level, the other two carry colour) makes it a plain “known channel”, and inpainting does the rest (Appendix I.3). No retraining is needed for either.
Why it matters today
Adding a classifier's slope to the score is the ancestor of the guidance used in every text-to-image model. The later refinement removes the separate classifier entirely: see the classifier-free guidance companion and the lesson's guided_noise.
6 Conclusion · original
“While our proposed sampling approaches improve results and enable more efficient sampling, they remain slower at sampling than GANs (Goodfellow et al. 2014) on the same datasets.”Song et al. (2020), §6
Everyday picture
The paper finishes with a map and an honest note. The map: every earlier score-based and diffusion method is a point on one landscape of SDEs, losses and solvers. The note: walking that landscape still takes hundreds of steps where a GAN takes one, and the many samplers bring many knobs to tune.
Why it matters today
Both open problems became research programmes: faster samplers and distillation into few-step students for the first; principled design spaces for the second. The rectified flow paper attacks the first by straightening the road until one step is enough.
Appendices, briefly · original
What they contain
- A, more general SDEs. The noise strength may be a matrix that depends on x; the reverse SDE and the probability flow gain one extra term each.
- B, the three SDEs derived. The limits of SMLD's and DDPM's chains as the number of steps grows, the variance formulas drawn above, and the bell-curve fades (equation 29).
- C, settings. VP uses β̄min = 0.1 and β̄max = 20 to match DDPM. Solvers stop at a small time ε rather than 0 (10−5 for VE; for VP, 10−3 when sampling and 10−5 for training and likelihoods), because VE's noise level jumps at t = 0 and VP's variance vanishes there.
- D, probability flow. The derivation uses the Fokker-Planck equation: fold the noise's spreading into the drift and the SDE becomes an ODE with the same crowd at every moment. Likelihoods use an adaptive Runge-Kutta solver and the trace estimator below. D.5 compares two differently built models' codes for the same image and finds them nearly identical.
- E and F, reverse diffusion and ancestral sampling. The reverse diffusion rule for any SDE, and DDPM-style ancestral steps adapted to SMLD models.
- G, predictor-corrector details. Algorithms 2 to 5, the step-size rule above, and a final denoising step using Tweedie's formula at the end of sampling, which the paper finds matters a lot for FID.
- H, architectures and training. 1.3 million iterations, an exponential moving average of the weights (0.999 for VE, 0.9999 for VP), and the 1024 × 1024 model.
- I, controllable generation. The identity used in §5, and a general recipe for inverse problems when a forward model p(y | x) is known.
The trace trick, decoded
Computing the divergence ∇·f̃ exactly means adding up the diagonal of a d × d matrix of slopes (the Jacobian), one network pass per dimension: 3,072 passes for a CIFAR-10 image. The paper instead uses the Skilling-Hutchinson estimator: squeeze the matrix between a random vector and itself.
In words: “the sum of the diagonal equals the average, over random vectors with mean 0 and identity covariance, of the vector times the matrix times the vector.” One backward pass computes εᵀ times the matrix, so each estimate costs about one network call, and the estimate is an unbiased estimator.
With the numbers: take the 2 × 2 matrix [[2, 1], [0, 3]], whose diagonal sums to 5. The four sign vectors (±1, ±1) give εᵀAε = 6, 4, 4 and 6; the average is exactly 5. The off-diagonal 1 appears with sign +1 half the time and −1 half the time, and cancels.
In Python:
# A: a 2 × 2 matrix of slopes whose diagonal sums to 2 + 3 = 5
A = [[2, 1], [0, 3]]
# εᵀ A ε for every random sign vector ε in {−1, +1}²
values = []
for e1 in (-1, 1):
for e2 in (-1, 1):
e = [e1, e2]
values.append(sum(e[i] * A[i][j] * e[j] for i in range(2) for j in range(2)))
values # → [6, 4, 4, 6]
# their average is the diagonal's sum exactly
sum(values) / len(values) # → 5.0
Two small slips in the source
In Appendix A's sliced score matching loss (equation 19) the second term is printed as v⊤sθv, a vector times a vector times a vector; for sliced score matching it has to be v⊤∇sθv, with the Jacobian of the score in the middle, exactly the shape of the trace trick above. And in Appendix I.4's equation 50, the integral is written over dy where the variable being integrated is y(t). Neither affects anything else in the paper.
What changed since 2020
The framework stuck; many of its specific choices moved on.
| In the paper | Common today | Why | Lesson |
|---|---|---|---|
| Predictor-corrector SDE sampling with 1,000 to 2,000 steps | Deterministic ODE solvers (DDIM and its successors) with tens of steps | Sampling cost is paid on every image | diffusion |
| The ODE's road inherited from a noising SDE | The road chosen directly, and often straight: flow matching and rectified flow | Straighter roads need fewer solver steps | diffusion |
| A separate noisy-image classifier for conditioning | Classifier-free guidance from one network | No second network to train, and it works with text | diffusion |
| A U-Net on 32 × 32 pixels | Transformers over patches of an autoencoder's latent | Cheaper per image, and they scale | diffusion |
| Predicting the score (or the noise) | Predicting the noise, the clean image, or a velocity, which are related by simple rescalings | Different targets balance the noise levels differently | diffusion |
Glossary
Every term with hover guidance on this page, in one place.