Denoising Diffusion Implicit Models, annotated
How to read this page
Nothing on this page assumes you already know the jargon. Three things help:
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation is followed by a table of its symbols.
- The diagrams and pictures are live: hover or tap the parts, drag the sliders, and the readouts say what changed.
Each idea is explained in the same order: an everyday picture, a tiny example you could check by hand, a diagram, then the math, and finally why it still matters.
One warning about notation, before anything else. In the DDPM paper and in the diffusion lesson, αt is the share of variance one step keeps and ᾱt (“alpha bar”) is the signal left after all steps so far. This paper drops the bar: its αt means DDPM's ᾱt, the signal share left at level t. (Appendix C.2 explains why: one list of numbers is all you need to choose, and it makes skipping steps easy to write.) This page follows the paper, so every α below is a running product, going from α0 = 1 (clean) down to nearly 0 (pure noise).
The one-pixel toy chain. A clean pixel x0 = 2.0 and two noise levels: α1 = 0.96 (lightly noised) and α2 = 0.64 (heavily noised). One noise draw ε = 0.5 takes the pixel to x2 = √0.64 × 2.0 + √0.36 × 0.5 = 0.8 × 2.0 + 0.6 × 0.5 = 1.9. The paper uses the same arithmetic with 1,000 levels and whole images; the diffusion lesson builds it from scratch in NumPy on dots in a plane.
Abstract
“To accelerate sampling, we present denoising diffusion implicit models (DDIMs), a more efficient class of iterative implicit probabilistic models with the same training procedure as DDPMs.”Song, Meng and Ermon (2020), Abstract. Read the original
Everyday picture
A diffusion model generates a picture like a restorer who removes dust one speck at a time, a thousand passes per picture. This paper keeps the same restorer, with the same training, and changes only the instructions: look at the dusty picture, imagine the fully clean result, then jump most of the way there at once. Twenty looks instead of a thousand.
What the paper claims
- The training loss of a DDPM is shared by a whole family of noising processes, most of which are not Markov chains. Any member of the family can be used for sampling, with the same trained network.
- One member samples with no randomness at all after the starting noise. It is an implicit generative model, named DDIM.
- DDIM samples of comparable quality come 10× to 50× faster in wall-clock time. It lets you trade steps for quality, interpolate meaningfully between samples by blending their starting noise, and encode a real image into noise and decode it back with very low error.
Why it matters today
This is the paper that made “sample in 20 to 50 steps” normal. Its update rule is also the template that classifier-free guidance plugs into (see the classifier-free guidance companion), and its reading of sampling as solving an ordinary differential equation leads straight to the straight-path samplers of the flow matching companion.
1 Introduction · original
“For example, it takes around 20 hours to sample 50k images of size 32 × 32 from a DDPM, but less than a minute to do so from a GAN on a Nvidia 2080 Ti GPU.”Song, Meng and Ermon (2020), §1
Everyday picture
A GAN paints a picture in one stroke: one pass through its network. A DDPM paints it in a thousand strokes, and each stroke is a full pass through a network of similar size. Diffusion models had caught up with GANs on quality without the fragile two-player training, but they paid for it at every single sample.
Tiny example
The paper's own numbers: 20 hours for 50,000 images of 32 × 32 pixels, and nearly 1,000 hours at 256 × 256. The first is 20 × 3,600 / 50,000 = 1.44 seconds per image, or 1.44 milliseconds per network call when every image needs 1,000 calls. The call is not slow; there are just a thousand of them.
In Python:
hours, images, steps = 20, 50_000, 1000
# seconds per image
per_image = hours * 3600 / images
round(per_image, 2) # → 1.44
# milliseconds per network call
round(per_image / steps * 1000, 2) # → 1.44
The paper's three findings
- Quality at speed: when sampling is accelerated 10× to 100×, DDIMs give better samples than DDPMs run with the same number of steps.
- Consistency: start from the same noise and sample with chains of different lengths, and the results share their high-level features. That is not true for DDPMs.
- Interpolation: because of that consistency, blending two starting noises gives a meaningful blend of the two pictures.
Why it matters
Sampling cost is paid on every picture anyone ever asks for; training is paid once. A trick that cuts sampling cost without touching training is worth more than almost any training improvement.
2 Background · original
Everyday picture
A DDPM has two halves. The forward process is fixed arithmetic that fades a picture into static a little at a time. The reverse process is a learned chain that runs the road backwards. Both are explained step by step in the DDPM companion; this section only restates them in this paper's notation.
Tiny example
In the toy chain, one forward step goes from level 1 (α1 = 0.96) to level 2 (α2 = 0.64). In DDPM's language, that step keeps α2/α1 = 2/3 of the variance and replaces 1/3 with noise. Or skip the chain and jump straight from the clean pixel: x2 = 0.8 × 2.0 + 0.6 × ε.
In words: “each forward step shrinks the previous value by the square root of the variance share it keeps and adds noise to fill the rest; and the whole chain collapses into one jump: the clean value, shrunk by √αt, plus noise scaled by √(1 − αt).”
With the numbers: the step from level 1 to level 2 keeps a share 0.64/0.96 = 0.6667 and adds noise of variance 0.3333. The jump: x2 = 0.8 × 2.0 + 0.6 × 0.5 = 1.9. These are equations 3 and 4 of the paper.
In Python:
import math
x0, eps = 2.0, 0.5
alpha1, alpha2 = 0.96, 0.64
# the step from level 1 to level 2: kept share α_t/α_(t−1) and noise variance 1 − α_t/α_(t−1)
round(alpha2 / alpha1, 4), round(1 - alpha2 / alpha1, 4) # → (0.6667, 0.3333)
# the jump: x_t = √α_t x_0 + √(1 − α_t) ε
x2 = math.sqrt(alpha2) * x0 + math.sqrt(1 - alpha2) * eps
round(x2, 4) # → 1.9
The training loss
The model is a set of noise guessers εθ(t), one per level (in practice one network told the level). Its bound, as in the DDPM paper, simplifies to a weighted squared error on the noise:
In words: “at every level, noise a real example in one jump, ask the network what noise was added, and score the squared miss; weight each level by γt and add them up.”
With the numbers: at level 2 the network sees 1.9 and guesses 0.4 instead of the true 0.5: a squared miss of 0.01. DDPM's best samples came from γ = 1, every level weighted the same (the paper calls it L1); this is the loss the next section is about.
In Python:
eps_true, eps_guess, gamma = 0.5, 0.4, 1.0
# γ_t ‖ε_θ − ε_t‖² for one example at one level
round(gamma * (eps_guess - eps_true) ** 2, 4) # → 0.01
Why it matters
The number of levels T is set before training, and in the DDPM recipe sampling must visit all of them: T = 1,000 levels means 1,000 network calls, one after another, with no way to run them in parallel. The rest of the paper asks whether that coupling is really necessary.
In code: add_noise is the one-jump formula and train_noise_predictor trains on the loss with γ = 1.
3 Variational inference for non-Markovian forward processes · original
“Our key observation is that the DDPM objective in the form of Lγ only depends on the marginals q(xt|x0), but not directly on the joint q(x1:T|x0).”Song, Meng and Ermon (2020), §3
Everyday picture
Picture a class of students graded only on single snapshots: “at minute 5, how noisy was the photo?”, “at minute 20?”. Nobody ever asks how the noise got from minute 5 to minute 20. Then many different noising stories, some fresh-noise-every-minute, some one-noise-turned-up-slowly, give identical exams, and the same student passes all of them. The training loss above is exactly such an exam: every term looks at one level at a time, through the one-jump formula.
Tiny example
Two recipes for noising the toy pixel to both levels:
- Recipe A (DDPM, fresh noise each step): x1 = √0.96 × 2 + √0.04 × ε1, then x2 = √(2/3) × x1 + √(1/3) × ε2, two independent draws.
- Recipe B (one noise, turned up): draw one ε and set x1 = √0.96 × 2 + 0.2ε and x2 = 0.8 × 2 + 0.6ε.
Recipe A gives x2 a centre of √(2/3) × √0.96 × 2 = 1.6 and a variance of (2/3) × 0.04 + 1/3 = 0.36. Recipe B gives 1.6 and 0.6² = 0.36. Same bell curve at every level, so the same training loss, yet the two recipes tell very different stories in between.
In Python:
import math
x0 = 2.0
# Recipe A: x_1 ~ N(√0.96 x0, 0.04), then x_2 = √(2/3) x_1 + √(1/3) ε_2
mean_A = math.sqrt(2 / 3) * math.sqrt(0.96) * x0
var_A = (2 / 3) * 0.04 + 1 / 3
round(mean_A, 4), round(var_A, 4) # → (1.6, 0.36)
# Recipe B: x_2 = 0.8 x0 + 0.6 ε, one ε for every level
round(0.8 * x0, 4), round(0.6 ** 2, 4) # → (1.6, 0.36)
Why it matters
If the loss cannot tell the recipes apart, then a network trained for recipe A (the only one anyone had trained) is equally a trained network for recipe B, and for everything in between. And each recipe comes with its own way of running backwards. Choosing the recipe becomes a sampling-time decision.
3.1 Non-Markovian forward processes · original
Everyday picture
Think of the noise level as a dimmer and the noise itself as a pattern of static. DDPM re-rolls the static every time the dimmer moves. The new family lets you keep some or all of the old pattern as the dimmer moves: a dial σt says how much fresh static to roll in at each step. Turn it to zero and the same static pattern just gets louder.
Tiny example
Suppose you know the clean pixel (2.0) and you see it at level 2 (1.9). Where was it at level 1? Work out the noise pattern it must carry: (1.9 − 0.8 × 2.0)/0.6 = 0.5. With σ = 0 the answer is certain: the same pattern at the lower volume, √0.96 × 2.0 + √0.04 × 0.5 = 2.0596. With DDPM's σ = 0.1925, the answer is a bell curve centred at 1.9868 with that spread.
In words: “one level back, the value is the clean value at the lower level's volume, plus the noise pattern read off xt at a slightly lower volume, plus fresh noise of spread σt; the two noise volumes are chosen so they add up to exactly the right total.”
With the numbers: the noise pattern is (1.9 − 1.6)/0.6 = 0.5. With σ2 = 0: centre 0.9798 × 2.0 + √0.04 × 0.5 = 1.9596 + 0.1 = 2.0596 and no spread. With σ2² = 0.0370 (DDPM's value, below): centre 1.9596 + √0.00296 × 0.5 = 1.9868, which is exactly the DDPM backward centre μ̃ of the DDPM companion's §2.
In Python:
import math
x0, x2 = 2.0, 1.9
alpha1, alpha2 = 0.96, 0.64
# the noise pattern x_t carries: (x_t − √α_t x_0) / √(1 − α_t)
pattern = (x2 - math.sqrt(alpha2) * x0) / math.sqrt(1 - alpha2)
round(pattern, 4) # → 0.5
def centre(sigma):
# √α_(t−1) x_0 + √(1 − α_(t−1) − σ_t²) · pattern
return math.sqrt(alpha1) * x0 + math.sqrt(1 - alpha1 - sigma ** 2) * pattern
sigma_ddpm = math.sqrt((1 - alpha1) / (1 - alpha2)) * math.sqrt(1 - alpha2 / alpha1)
round(centre(0.0), 4), round(centre(sigma_ddpm), 4) # → (2.0596, 1.9868)
The paper builds the whole forward process backwards from these pieces (its equation 6): first xT from the one-jump formula, then each xt−1 from xt and x0. The centre was chosen so that every level still has the right one-jump bell curve, N(√αt x0, (1 − αt)I) (Lemma 1, checked with numbers in the appendix). Because xt−1 depends on x0 as well as xt, the process is not a Markov chain any more; the paper calls it non-Markovian.
Hover or tap a circle. Start with x0, the clean picture, at the bottom of either side.
Reading it: both sides show how the noisy levels x1, x2, x3 are made from a clean picture x0 (shaded, because it is observed). On the left, DDPM's chain: each level is made from the one just below it and nothing else, so the arrows climb one rung at a time. On the right, the non-Markovian process: the top level x3 is made straight from x0 (the long dashed arrow), and then the process works downwards, each level made from the one above it and from x0 (the other dashed arrows). The pictures look different, but hover any level and the note says the same thing on both sides: its bell curve given x0 is identical. That shared bell curve is all the training loss ever sees.
See it: the same bell curves, different journeys
Try it: drag η from 1 (DDPM's fresh noise at every level) down to 0 (one noise pattern, turned up). Watch the individual paths, then watch the shaded band and the readout.
Reading it: the x-axis is the level t, from the clean pixel (x0 = 1, on the left) to pure noise (t = 1,000, on the right), on the paper's schedule. Each thin line (12 are drawn) is one run of the forward process of equation 7, built from the top level down as the paper defines it, with σt = η × DDPM's value. The shaded band is the one-jump bell curve's centre ± 2 standard deviations at each level: what the training loss sees. At η = 1 the paths zigzag, because every level rolls fresh noise. At η = 0 each path is a smooth curve, because one noise pattern is simply turned up. The readout measures, over 2,000 runs, how spread out the paths are at level 500: it matches the band's √(1 − α500) = 0.960 at every η, up to sampling wobble in the second decimal place. Computed live, in your browser.
Why it matters today
The whole paper rests on this freedom. Since the network was only ever asked about one level at a time, you may choose the in-between story after training, and the story with no fresh noise turns out to be the most useful.
3.2 Generative process and unified variational inference objective · original
Everyday picture
Equation 7 needs the clean picture, which a sampler does not have. So the sampler first guesses the clean picture from what it sees, then uses equation 7 as if the guess were true. A detective who cannot see the culprit sketches a best guess from the clues, then reasons from the sketch.
Tiny example
The network looks at x2 = 1.9 and guesses the noise as 0.5 (perfect). Undo the one-jump formula: (1.9 − 0.6 × 0.5)/0.8 = 2.0, the true clean pixel. A slightly wrong guess, 0.4, gives (1.9 − 0.24)/0.8 = 2.075.
In words: “the predicted clean picture is the noisy one with the guessed noise removed, scaled back up by the signal it lost.”
With the numbers: (1.9 − 0.6 × 0.5)/0.8 = 2.0; with the guess 0.4, (1.9 − 0.6 × 0.4)/0.8 = 2.075.
In Python:
import math
x_t, alpha_t = 1.9, 0.64
# f_θ(x_t) = (x_t − √(1 − α_t) ε_θ) / √α_t, for a perfect and a slightly wrong guess
[round((x_t - math.sqrt(1 - alpha_t) * g) / math.sqrt(alpha_t), 4) for g in (0.5, 0.4)] # → [2.0, 2.075]
The generative process then starts at pure noise, pθ(xT) = N(0, I), and uses equation 7 with the guess in place of x0:
In words: “to step back, guess the clean picture and take the forward process's own backward step as if that guess were right; at the very last step, just output the guess (with a little noise so the model gives every picture some probability).”
With the numbers: from x2 = 1.9 with the perfect guess, the step to level 1 is qσ(x1 | 1.9, 2.0): centre 2.0596 with σ = 0, or 1.9868 with DDPM's σ, the same numbers as in §3.1. With the guess 0.4 (so f = 2.075) and DDPM's σ, the centre moves to 2.0548.
In Python:
import math
alpha1, alpha2, x2 = 0.96, 0.64, 1.9
sigma = math.sqrt((1 - alpha1) / (1 - alpha2)) * math.sqrt(1 - alpha2 / alpha1)
def step_centre(f, sigma):
# the centre of q_σ(x_1 | x_2, f): eq 7 with x_0 replaced by the guess f
pattern = (x2 - math.sqrt(alpha2) * f) / math.sqrt(1 - alpha2)
return math.sqrt(alpha1) * f + math.sqrt(1 - alpha1 - sigma ** 2) * pattern
round(step_centre(2.0, 0.0), 4), round(step_centre(2.0, sigma), 4) # → (2.0596, 1.9868)
round(step_centre(2.075, sigma), 4) # → 2.0548
Theorem 1: one training loss for the whole family
Each σ gives a different forward process and a different generative process, so on paper each has its own variational bound Jσ, and it looks as if each would need its own training run. Theorem 1 says no:
In words: “whatever amount of fresh noise you pick, the proper training objective for that choice is the familiar noise-guessing loss with some positive weight on each level, plus a number that does not depend on the network.”
With the numbers: in the toy, at level 2 with DDPM's σ, the bound's term is a KL divergence between two bell curves of equal width whose centres differ only because the guess differs. Guess 0.4 gives a KL of 0.0625; guess 0.3 gives 0.25. Divide by the squared misses (0.01 and 0.04): both give the same weight, γ2 = 6.25. The bound really is the noise loss, reweighted.
In Python:
import math
x0, x2, eps = 2.0, 1.9, 0.5
alpha1, alpha2 = 0.96, 0.64
sigma = math.sqrt((1 - alpha1) / (1 - alpha2)) * math.sqrt(1 - alpha2 / alpha1)
def centre(f):
# the centre of q_σ(x_1 | x_2, f)
return math.sqrt(alpha1) * f + math.sqrt(1 - alpha1 - sigma ** 2) * (x2 - math.sqrt(alpha2) * f) / math.sqrt(1 - alpha2)
for guess in (0.4, 0.3):
f = (x2 - math.sqrt(1 - alpha2) * guess) / math.sqrt(alpha2)
# KL between two bell curves of equal variance σ²: (difference of centres)² / (2σ²)
kl = (centre(x0) - centre(f)) ** 2 / (2 * sigma ** 2)
print(round(kl, 4), round(kl / (eps - guess) ** 2, 4)) # → 0.0625 6.25 0.25 6.25
Here is the key step. If every level had its own separate network, the best network for Lγ would not depend on γ at all: each level's term is minimised on its own, and scaling a term does not move its minimum. So the weights do not matter, the network trained with γ = 1 is optimal for every Jσ, and DDPM's trained networks can be used for any member of the family. (In practice one network is shared across levels, so this is an argument rather than a guarantee; the experiments are the evidence that it holds.)
One detail for careful readers: the constant printed in the paper's proof (Appendix B, equation 35) is 1/(2dσt²αt), which for the toy would be 21.09, not the 6.25 computed directly above. The proof drops the factor that turns a miss in x0 into a miss in ε and the centre's coefficient on x0. Only the existence of some positive weight matters for the theorem, and that survives.
Why it matters today
Theorem 1 is why a single diffusion checkpoint can be paired with many samplers, chosen by whoever runs it. Every “sampler” menu in an image tool rests on this separation between how a model was trained and how it is sampled.
4 Sampling from generalized generative processes · original
Everyday picture
Same chef, same training, a new recipe card. This section writes the sampling step for any σ, picks out two special cards (DDPM and DDIM), and then shows how to skip most of the steps.
4.1 Denoising diffusion implicit models · original
“Different choices of σ values results in different generative processes, all while using the same model εθ, so re-training the model is unnecessary.”Song, Meng and Ermon (2020), §4.1
Everyday picture
One step back is three moves. Predict the finished picture. Point back along the noise you believe is in the current one, at the volume the next level should have. Then, optionally, add a pinch of fresh noise. DDPM always adds the pinch; DDIM never does.
Tiny example
From x2 = 1.9 with a noise guess of 0.5: the predicted clean pixel is 2.0; the next level wants √0.96 × 2.0 = 1.9596 of signal and a noise volume of √0.04 = 0.2. With no fresh noise, x1 = 1.9596 + 0.2 × 0.5 = 2.0596. With DDPM's σ = 0.1925 and a fresh draw of −1.0: 1.9596 + 0.0544 × 0.5 + 0.1925 × (−1.0) = 1.7944.
In words: “the next, cleaner value is the predicted clean picture at the next level's volume, plus the guessed noise at the next level's volume (less whatever is handed to fresh noise), plus the fresh noise.”
With the numbers: σ = 0: 0.9798 × 2.0 + 0.2 × 0.5 + 0 = 2.0596. σ = 0.1925, fresh εt = −1.0 (illustrative): 1.9596 + 0.0272 − 0.1925 = 1.7944. This is equation 12, the paper's general sampler; α0 is defined as 1, so the last step outputs the predicted clean picture.
In Python:
import math
x_t, eps_hat = 1.9, 0.5
alpha_t, alpha_prev = 0.64, 0.96
def step(sigma, fresh):
x0_pred = (x_t - math.sqrt(1 - alpha_t) * eps_hat) / math.sqrt(alpha_t)
toward_xt = math.sqrt(1 - alpha_prev - sigma ** 2) * eps_hat
return math.sqrt(alpha_prev) * x0_pred + toward_xt + sigma * fresh
sigma_ddpm = math.sqrt((1 - alpha_prev) / (1 - alpha_t)) * math.sqrt(1 - alpha_t / alpha_prev)
# DDIM (σ = 0), then DDPM with an illustrative fresh draw of −1.0
round(step(0.0, 0.0), 4), round(step(sigma_ddpm, -1.0), 4) # → (2.0596, 1.7944)
Hover or tap a box. One network call per step; everything else is arithmetic.
Reading it: the only expensive box is the network at the top, run once per step. Its noise guess is used twice: on the left to predict the clean picture, on the right as the direction pointing back toward the current noisy input. The two meet at the plus sign, each scaled to the next level's volume, and the sum leaves at the bottom right as xt−1. The box at the bottom left, fresh noise, is the only source of randomness: it is on for DDPM and switched off (σt = 0) for DDIM. With it off, the step is a fixed function of xt.
The two special cases
One particular σ makes the forward process Markovian again, and the sampler becomes exactly DDPM:
In words: “DDPM's fresh noise has the spread of one forward step's noise, shrunk by the ratio of the two levels' noise.”
With the numbers: √(0.04/0.36) × √(1 − 0.64/0.96) = 0.3333 × 0.5774 = 0.1925, so σ2² = 0.0370: DDPM's backward variance β̃ from the DDPM companion, written in this paper's letters.
In Python:
import math
alpha_prev, alpha_t = 0.96, 0.64
sigma = math.sqrt((1 - alpha_prev) / (1 - alpha_t)) * math.sqrt(1 - alpha_t / alpha_prev)
round(sigma, 4), round(sigma ** 2, 4) # → (0.1925, 0.037)
# DDPM's β̃_t = (1 − ᾱ_(t−1))/(1 − ᾱ_t) · β_t, with ᾱ = this paper's α and β_t = 1 − α_t/α_(t−1)
round((1 - alpha_prev) / (1 - alpha_t) * (1 - alpha_t / alpha_prev), 4) # → 0.037
The other special case is σt = 0 for every t. The random term vanishes, and after the starting noise xT every step is fixed arithmetic: the same xT always gives the same picture. A model that makes samples by pushing a random starting code through a fixed procedure is an implicit generative model, like a GAN's generator. The authors name this one the denoising diffusion implicit model, DDIM, “because it is an implicit probabilistic model trained with the DDPM objective”. (Strictly, Theorem 1 needs σ > 0; a footnote notes σ = 0 is the limit of ever smaller σ.)
Why it matters today
Equation 12 is still how most diffusion code writes a sampling step: predict x0, re-noise toward the next level, optionally add noise. The σ = 0 case is the one people usually mean by “the DDIM sampler”.
In code: ddim_step is equation 12 with σ = 0 (the lesson writes it as “predict x̂0, then re-noise”), and ddpm_step is DDPM's small step.
4.2 Accelerated generation processes · original
“In principle, this means that we can train a model with an arbitrary number of forward steps but only sample from some of them in the generative process.”Song, Meng and Ermon (2020), §4.2
Everyday picture
A staircase of 1,000 steps, but the exam only ever asked about individual steps, never about climbing them in order. So build a forward process that visits only every 100th step, check that each visited step has the right bell curve, and walk down that short staircase instead. The network already knows every step it will be asked about.
Tiny example
With T = 3 levels and the sub-sequence τ = [1, 3] (the paper always ends τ at T), sampling goes x3 → x1 → x0: two network calls instead of three. The step from 3 to 1 is equation 12 with t = 3 and t − 1 replaced by 1: the formula never needed the two levels to be neighbours.
Hover or tap a circle. Sampling reads left to right.
Reading it: sampling runs left to right along the solid arrows: start at x3, jump straight to x1, then to the picture x0. Level x2 (dashed) is never visited. In the paper's construction it hangs off x0 alone, like a star's spoke, so it still appears in the training bound but plays no part in sampling. Each solid arrow is one network call; the number of arrows, not T, is the cost of a sample.
Choosing which levels to visit
Appendix D.2 gives two rules for choosing S levels out of T:
In words: “visit evenly spaced levels, or levels that bunch up near the clean end, where fine detail is decided; pick c so the last visited level lands near T.”
With the numbers: T = 1,000 and S = 10. Linear with c = 100: 100, 200, …, 1000. Quadratic with c = 10: 10, 40, 90, 160, 250, 360, 490, 640, 810, 1000 (these values of c are this page's choice). The paper used quadratic for CIFAR10 and linear elsewhere.
In Python:
import math
T, S = 1000, 10
# τ_i = ⌊c i⌋ with c = T/S, and τ_i = ⌊c i²⌋ with c = T/S²
[math.floor(T / S * i) for i in range(1, S + 1)] # → [100, 200, 300, 400, 500, 600, 700, 800, 900, 1000]
[math.floor(T / S ** 2 * i * i) for i in range(1, S + 1)] # → [10, 40, 90, 160, 250, 360, 490, 640, 810, 1000]
Why it matters today
Training with many fine levels and sampling a few coarse ones is now the norm. The paper also notes the training levels could even become continuous, which later diffusion work adopted (the classifier-free guidance paper trains in continuous time).
In code: ddim_path picks evenly spaced levels and jumps between them; ddim_sample keeps only the end point.
4.3 Relevance to neural ODEs · original
Everyday picture
With no fresh noise, sampling is a journey along a fixed road: at every point the network says which way to go, as a wind map would. Following such a map is solving an ordinary differential equation, and each DDIM jump is one Euler step: move in a straight line at the current direction, then look again. The trick is which “clock” you step along.
Tiny example
Rescale the toy pixel by its signal: x̄ = x/√α. At level 2 that is 1.9/0.8 = 2.375, and the noise-to-signal ratio σ = √((1 − α)/α) is 0.75. At level 1, σ = √(0.04/0.96) = 0.2041. In these units the DDIM jump is a straight line: x̄ changes by (0.2041 − 0.75) × 0.5 = −0.273, to 2.1021, which is 2.0596 after scaling back by √0.96. The same answer as equation 12.
In words: “measured in signal-rescaled units, one DDIM jump moves the point by the noise guess times the change in the noise-to-signal ratio.”
With the numbers: 1.9/0.8 + (0.2041 − 0.75) × 0.5 = 2.375 − 0.2730 = 2.1021; times √0.96 = 2.0596.
In Python:
import math
x_t, eps_hat = 1.9, 0.5
alpha_t, alpha_next = 0.64, 0.96
def ratio(a):
# σ = √((1 − α)/α): noise-to-signal
return math.sqrt((1 - a) / a)
round(ratio(alpha_t), 4), round(ratio(alpha_next), 4) # → (0.75, 0.2041)
xbar_next = x_t / math.sqrt(alpha_t) + (ratio(alpha_next) - ratio(alpha_t)) * eps_hat
round(xbar_next, 4), round(xbar_next * math.sqrt(alpha_next), 4) # → (2.1021, 2.0596)
Let the jumps shrink to nothing and equation 13 becomes an ODE, with σ(t) = √((1 − α)/α) as the clock and x̄ = x/√α as the position:
In words: “as the noise-to-signal ratio changes by a tiny amount, the rescaled point moves by the network's noise guess times that amount.”
With the numbers: at x̄ = 2.375, σ = 0.75, the network is asked about x̄/√(0.75² + 1) = 2.375/1.25 = 1.9 (the ordinary xt), guesses 0.5, so x̄ changes 0.5 times as fast as σ does. Run it backwards in σ and you generate; run it forwards and you encode a picture into its noise, which DDPM cannot do.
In Python:
import math
xbar, sigma = 2.375, 0.75
# the network is asked about x = x̄ / √(σ² + 1), which is x_t
round(xbar / math.sqrt(sigma ** 2 + 1), 4) # → 1.9
# and α = 1 / (1 + σ²) recovers the signal share
round(1 / (1 + sigma ** 2), 4) # → 0.64
Proposition 1: the same ODE as the probability flow, stepped on a different clock
Concurrent work (Song et al., on stochastic differential equations) had described a probability flow ODE. Proposition 1 says that, with a perfect network, the DDIM ODE is the same ODE as their “variance exploding” one. The difference is the Euler step: their natural step moves along time t, which in the same units reads
In words: “the same move, but the change in noise-to-signal ratio is estimated from the change in its square, halved and divided by the current ratio: correct for tiny steps, off for big ones.”
With the numbers: ½ × (0.0417 − 0.5625) / 0.75 × 0.5 = −0.1736, so x̄ = 2.2014 and x = 2.1569, not 2.0596. For a small jump the two agree; for this big one they do not.
In Python:
import math
x_t, eps_hat = 1.9, 0.5
alpha_t, alpha_next = 0.64, 0.96
s2_now, s2_next = (1 - alpha_t) / alpha_t, (1 - alpha_next) / alpha_next
xbar = x_t / math.sqrt(alpha_t) + 0.5 * (s2_next - s2_now) / math.sqrt(s2_now) * eps_hat
round(xbar, 4), round(xbar * math.sqrt(alpha_next), 4) # → (2.2014, 2.1569)
Try it: the toy data are two clean values, −2 and +2. Pick where to start, then read how far one big jump of each kind lands from the exact ODE solution, as the target level gets cleaner (to the right).
Hover or use the arrow keys to compare the three landing points at each target level.
Reading it: the x-axis is the target level's signal share α (to the right is cleaner); the y-axis is where one jump lands. The solid line is the exact answer, found by following the ODE in many tiny steps. Start at 1.9: nearly all the probability says the clean value is +2, so the network's noise guess barely changes along the road, and the DDIM jump (dashed) lies exactly on the exact line, at any distance. The time-stepped Euler jump (dotted) drifts away as the jump grows: 2.1569 against 2.0596 at α = 0.96. Now start at 0.5 at α = 0.2, where both clean values are plausible: the guess changes as the road bends toward +2, and both single jumps fall short. That is the error the paper's discussion hopes multi-step ODE methods can reduce. Computed live from equations 13 and 15.
Why it matters today
Reading a sampler as an ODE solver opened a whole line of work: better solvers take fewer steps for the same accuracy. It also makes diffusion models invertible, which image-editing methods use to find the noise that produces a given photo. Flow matching takes the idea further by training the road to be straight; see the flow matching companion.
In code: euler_step is the plain Euler step the lesson uses for flow matching, and ddim_step is equation 13 written in ordinary units.
5 Experiments · original
Everyday picture
Nothing is retrained. The experiments take networks already trained the DDPM way (T = 1,000, γ = 1) and change only two knobs at sampling time: how many levels to visit, S, and how much fresh noise to add, η.
Tiny example
The knob η scales DDPM's σ: η = 1 is DDPM, η = 0 is DDIM, and η = 0.5 adds half DDPM's fresh noise. For the toy jump from level 2 to level 1, that is 0.1925, 0 and 0.0962.
In words: “η is a dial from no fresh noise (0) to DDPM's amount (1); σ̂ is a larger amount DDPM's original code used for CIFAR10, the full noise of one forward step.”
With the numbers: η = 0.5 gives 0.0962. And σ̂ = √(1 − 0.64/0.96) = 0.5774, so σ̂² = 0.333: more than eight times the 0.04 of noise variance the destination level should hold in total. Over one tiny step the two amounts are close; over a big jump σ̂ swamps the picture with noise.
In Python:
import math
alpha_prev, alpha_t = 0.96, 0.64
def sigma(eta):
return eta * math.sqrt((1 - alpha_prev) / (1 - alpha_t)) * math.sqrt(1 - alpha_t / alpha_prev)
[round(sigma(eta), 4) for eta in (1.0, 0.0, 0.5)] # → [0.1925, 0.0, 0.0962]
sigma_hat = math.sqrt(1 - alpha_t / alpha_prev)
# σ̂², and how many times the level's whole noise variance 1 − α_(t−1) it is
round(sigma_hat, 4), round(sigma_hat ** 2 / (1 - alpha_prev), 2) # → (0.5774, 8.33)
The setup
| Setting | Value |
|---|---|
| Datasets | CIFAR10 (32 × 32), CelebA (64 × 64), LSUN Bedroom and Church (256 × 256) |
| Networks | DDPM's pretrained U-Net checkpoints for CIFAR10, Bedroom and Church; a new one trained on CelebA with the same L1 loss |
| Training levels | T = 1,000, schedule as in DDPM |
| Knobs | S = dim(τ), the number of levels visited; η, the fresh noise |
| Level choice | quadratic τ for CIFAR10, linear for the others |
| Quality measure | FID, lower is better |
Why it matters
Because the network is fixed, every difference in the results below is a difference between samplers, not between models.
5.1 Sample quality and efficiency · original
“Notably, DDIM is able to produce samples with quality comparable to 1000 step models within 20 to 100 steps, which is a 10× to 50× speed up compared to the original DDPM.”Song, Meng and Ermon (2020), §5.1
Everyday picture
A painter asked to finish in ten strokes instead of a thousand: the careful one who commits to each stroke does fine; the one who adds a random splash after every stroke, as if there were still hundreds to come, leaves a mess.
Tiny example
| Sampler | S = 10 | S = 100 | S = 1000 |
|---|---|---|---|
| DDIM (η = 0) | 13.36 | 4.16 | 4.04 |
| DDPM (η = 1) | 41.07 | 5.78 | 4.73 |
| DDPM (σ̂) | 367.43 | 9.99 | 3.17 |
At 10 steps, DDIM's FID is about a third of DDPM's (13.36 against 41.07). At the full 1,000 steps the order flips slightly: σ̂ is best (3.17), which the paper notes. The σ̂ row shows the danger computed above: at 10 steps it scores 367.43, with visible noise left in the images.
Hover or use the arrow keys to read the FID at each step count.
Reading it: the x-axis is the number of sampling steps S on a logarithmic scale (10, 20, 50, 100, 1,000); the y-axis is FID, lower is better. Each line is one sampler on one dataset, drawn from the η = 0 and η = 1 rows of the paper's Table 1 (σ̂ is left out: its 10-step values, 367 and 300, would flatten everything else). Follow any pair of lines from the right: at 1,000 steps DDIM and DDPM are close. Move left and the DDPM lines climb steeply while the DDIM lines climb gently, so the gap opens exactly where sampling is cheap. On CelebA, 20 DDIM steps (13.73) roughly match 100 DDPM steps (13.93), the comparison the paper draws.
Figure 4 of the paper adds that sampling time grows in a straight line with S. So if 1,000 steps take the 20 hours of §1, then 20 steps take about 20 × 20/1,000 = 0.4 hours and 100 steps about 2: the 10× to 50× of the abstract.
In Python:
hours_1000 = 20
# time is linear in the number of steps (Figure 4)
[round(hours_1000 * S / 1000, 1) for S in (20, 50, 100)] # → [0.4, 1.0, 2.0]
# the speed-ups
[1000 // S for S in (20, 50, 100)] # → [50, 20, 10]
Try it: the sampler on four blobs
The toy data are four tight blobs, like the diffusion lesson's but drawn tighter so stray samples stand out. Instead of a trained network, this page computes the perfect noise guess exactly (possible because the data are a known mixture of bell curves), then runs equation 12 on the paper's 1,000-level schedule. Try it: set S = 10 and compare η = 0 with η = 1, then switch on σ̂. Then drop S to 1, 2 and 3 and watch where the samples go.
Reading it: the grey dots are the data, four blobs around the centre; the coloured dots are 400 samples, each started from the same noise every time. At S = 1 every sample lands in the empty middle: from pure noise, the best guess of the clean picture is the average of all the data. By S = 3 the samples have split toward the four blobs; by S = 10 most sit on them, with a few stragglers in the gaps, and by S = 20 nearly all do. With a perfect noise guess on this small toy, η makes little difference (the readout's distance moves only slightly); the large η gap in Table 1 is a measurement on real images with a real, imperfect network, and the paper reports it rather than explaining it. The σ̂ button reproduces the other finding clearly: at 5 to 20 steps it leaves many samples scattered outside the blobs (about half of them at S = 10), the too-much-noise problem worked out above, and only at hundreds of steps does it catch up. The readout's distance averages two checks (each sample to its nearest data point, and each data point to its nearest sample), so collapsing onto one blob cannot score well.
Why it matters today
This table is where “20 to 50 steps” as a default came from. It also showed that fresh noise, necessary for many small steps, becomes harmful for a few big ones, a lesson later samplers kept.
In code: step_sweep runs the lesson's trained networks at several step counts and scores them with two_way_distance, the same two-way check as the readout above.
5.2 Sample consistency in DDIMs · original
“Interestingly, for the generated images with the same initial xT, most high-level features are similar, regardless of the generative trajectory.”Song, Meng and Ermon (2020), §5.2
Everyday picture
Give the same rough sketch to a painter in a hurry and to a painter with all day. Their finished pictures differ in detail but show the same scene, because the sketch already decided it. With DDIM, the starting noise xT is that sketch; with DDPM, the fresh noise added along the way keeps redrawing it.
Tiny example
In the toy below, 500 starting noises are each sampled twice, with S = 10 and with S = 100. With η = 0, every one of them lands on the same blob both times, and the two end points are on average 0.12 apart. With η = 1, only about a quarter land on the same blob: no better than chance with four blobs.
Reading it: each hollow ring is where one starting noise ends after 10 steps, each filled dot where the same starting noise ends after 100 steps, and a line joins the pair (80 pairs are drawn; the readout counts 500). With DDIM most lines are short, and even the longer ones (a 10-step sample caught in a gap) end on the blob that sample was already heading for: the short and the long run agree, and the long run only tidies the detail. Switch to DDPM and the lines stretch across the picture, from blob to blob: the starting noise decided almost nothing, because the fresh noise added on the way redecided it.
The paper's conclusion is that xT alone works as a latent code for the picture: high-level features are fixed by xT, and more steps only refine details.
Why it matters today
Consistency is what makes a seed meaningful. Re-running an image generator with the same seed and a different step count gives a recognisably similar picture because modern deterministic samplers behave like DDIM.
5.3 Interpolation in deterministic generative processes · original
Everyday picture
If each starting noise is a sketch, you can blend two sketches and see what the painter makes of the blend. With a GAN this gives a smooth morph between two faces; the paper shows DDIM does the same, which DDPM cannot (its same starting noise gives many different pictures).
Tiny example
There is a catch in how to blend. Two unit-length noise directions, (1, 0) and (0, 1), averaged in a straight line give (0.5, 0.5), of length only 0.707: a weak, atypically quiet noise. Blend along the arc instead and the midpoint is (0.707, 0.707), still of length 1. In high dimensions this matters: a noise vector with 3,072 numbers (a CIFAR10 image) almost always has length close to √3,072 ≈ 55.4, so a straight-line midpoint of two of them, about 39.2 long, looks nothing like real starting noise.
In words: “measure the angle between the two noises, then walk along the arc between them, a fraction α of the way; this is spherical linear interpolation (slerp).”
With the numbers: (1, 0) and (0, 1) have cosine 0, so θ = π/2 and sin θ = 1. At α = 0.5 both weights are sin(π/4) = 0.7071, giving (0.7071, 0.7071), length 1. (Here α is the blend fraction; the paper reuses the letter.)
In Python:
import math
a, b = [1.0, 0.0], [0.0, 1.0]
cos = sum(x * y for x, y in zip(a, b)) / (math.hypot(*a) * math.hypot(*b))
theta = math.acos(cos)
alpha = 0.5
w_a = math.sin((1 - alpha) * theta) / math.sin(theta)
w_b = math.sin(alpha * theta) / math.sin(theta)
mid = [w_a * x + w_b * y for x, y in zip(a, b)]
[round(v, 4) for v in mid], round(math.hypot(*mid), 4) # → ([0.7071, 0.7071], 1.0)
# a straight-line midpoint is shorter: length 0.7071 of each end, e.g. √3072 → about 39.2
round(math.sqrt(3072), 1), round(0.7071 * math.sqrt(3072), 1) # → (55.4, 39.2)
Try it: two starting noises that DDIM sends to different blobs; slide α to blend them along the arc.
Reading it: the grey dots are the data. The dashed arc across the top is the path of blended starting noises, from the first noise (square, on the right) to the second (triangle, on the left); the ring marks the current blend. Noises and samples are both two numbers here, so they share one picture. The solid curve is where DDIM with 50 steps sends every point of that arc, and the large dot is where the current blend goes. As α grows the output first stays inside the first blob, then crosses quickly to the next blob and settles there: the map from noise to sample is continuous, but it hurries across the empty gaps between blobs. On real images those gaps are much smaller than in this toy, which is why the paper's interpolations look like smooth morphs.
Why it matters today
Walking in noise space, and editing a picture by finding its noise and changing the prompt, both depend on the deterministic map from xT to a picture that DDIM introduced.
5.4 Reconstruction from latent space · original
Everyday picture
If you can drive a road from noise to picture, you can drive it back from picture to noise. Encode a real photo by running the ODE the other way, decode the noise you get, and see how close you come back.
Tiny example
In the toy chain, encoding the clean pixel 2.0 to level 2 with a perfect noise guess (0.5) lands on 1.9; decoding 1.9 with the same guess returns 2.0 exactly. Errors in a real network, and the size of each jump, are what keep the round trip from being perfect.
| Steps S | 10 | 100 | 1000 |
|---|---|---|---|
| Error (per-dimension squared error, pixels scaled to [0, 1]) | 0.014 | 0.0009 | 0.0001 |
Hover or use the arrow keys to read the error at each step count.
Reading it: the x-axis is S, the steps used both to encode and to decode, from 10 to 1,000 on a logarithmic scale; the y-axis is the mean squared round-trip error per number. The solid line is computed live on this page's four-blob toy with a perfect noise guess: 200 data points are encoded to noise and decoded back. The dashed line is the paper's Table 2 for CIFAR10 with its trained network. Both fall steadily with S and agree closely up to 50 steps (0.013 against 0.014 at 10). Beyond that the paper's error levels off near 0.0001, while the toy's, with its perfect noise guess, keeps falling. The error comes from taking big Euler steps, not from any randomness, since there is none.
Why it matters today
This “inversion” is the basis of many diffusion editing methods: invert a real photo to its noise, then regenerate it with a changed prompt or a guided sampler. DDPM, with its fresh noise, cannot do it.
6 Related work · original
Everyday picture
DDPMs and score networks (NCSN) had arrived at the same place from two directions: both learn to denoise at many noise levels and both sample by something like Langevin dynamics, many small, noisy steps. DDIM steps outside that frame.
What it says
- Langevin dynamics discretizes a gradient flow, which the paper links to both families needing many steps.
- DDIM is an implicit model, where samples are fixed functions of the latent noise; that makes it resemble GANs and invertible flows, including meaningful interpolation.
- It is derived purely variationally, so the restrictions of Langevin dynamics do not apply; the authors suggest this may partly explain the better few-step quality.
Tiny example
A Langevin step on the toy pixel: move a little along the score and add fresh noise, so two runs from 1.9 separate after one step. A DDIM step from 1.9 gives 2.0596 every time. Same network, two kinds of sampler.
Why it matters today
The split between stochastic samplers (Langevin-like, DDPM-like) and deterministic ones (ODE-like, DDIM-like) is still how samplers are classified. See the DDPM companion's related work for the score-matching side.
7 Discussion · original
“…it would be interesting to see if methods that decrease the discretization error in ODEs, including multi-step methods such as Adams-Bashforth…, could be helpful for further improving sample quality in fewer steps.”Song, Meng and Ermon (2020), §7
What the paper concludes
- DDIMs sample much faster than DDPMs and NCSNs and interpolate meaningfully in their latent space.
- The non-Markovian view suggests noising processes that are not Gaussian at all; Appendix A sketches one for categories (discrete data).
- Since DDIM sampling is an ODE solver, better ODE solvers may need fewer steps still.
Tiny example
A multi-step method uses the previous step's direction as well as the current one to predict how the road bends. In the ODE figure above, starting at 0.5, one jump misses because the direction changes along the way: a method that estimates that change from two recent guesses can correct for it without extra network calls.
Why it matters today
That suggestion was taken up: a family of later samplers applies higher-order and multi-step ODE methods to exactly this equation, reaching good quality in 10 to 20 steps.
Appendices A to D · original
A: a discrete case
For data that are categories (a one-hot vector with K options), the paper defines a noising process that mixes the true category with a uniform choice, and a backward step that picks the current value, the predicted clean value, or a uniform draw with set probabilities. As σt shrinks the step becomes less random. The bound's terms turn into ordinary classification losses. The paper leaves experiments to future work.
B: Lemma 1, checked with numbers
Lemma 1 says equation 7 keeps every level's bell curve right. The centre's check is immediate (the noise pattern averages to 0). The variance check is one line:
In words: “the fresh noise contributes σt², and the noise carried over from xt contributes the rest, so together they give exactly the variance level t − 1 should have, whatever σt is.”
With the numbers: σ2² = 0.0370: 0.0370 + (0.04 − 0.0370)/0.36 × 0.36 = 0.04. With σ2 = 0: 0 + 0.04 = 0.04. Both equal 1 − α1.
In Python:
import math
alpha_prev, alpha_t = 0.96, 0.64
sigma_ddpm = math.sqrt((1 - alpha_prev) / (1 - alpha_t)) * math.sqrt(1 - alpha_t / alpha_prev)
for sigma in (sigma_ddpm, 0.0):
total = sigma ** 2 + (1 - alpha_prev - sigma ** 2) / (1 - alpha_t) * (1 - alpha_t)
print(round(total, 4)) # → 0.04 0.04
Appendix B also proves Theorem 1 (see the note in §3.2 about its printed constant) and Proposition 1, by rewriting both ODEs in the rescaled units x̄ and σ and finding the same equation. A few equations in the appendices carry small typographical slips that do not affect the argument: equation 46 writes x(t) = √α(t) x(t) where x(0) is meant, equation 48 likewise, equation 56 feeds xτi−1 to the network where xτi is meant, and equation 62 writes the variance as σtI rather than σt²I.
C: accelerated processes and the notation
Appendix C.1 builds the short forward process of §4.2 formally. Appendix C.2 translates to DDPM's letters: this paper's αt is DDPM's ᾱt, so DDPM's per-step values are ratios of neighbours:
In words: “a DDPM step's kept share is the ratio of this paper's signal shares at two neighbouring levels, and its noise dose is what is left.”
With the numbers: the DDPM companion's toy chain has β = 0.1 then 0.2, so this paper's α list is 1, 0.9, 0.72. Ratios of neighbours give back 0.9 and 0.8, and β = 0.1 and 0.2.
In Python:
alphas = [1.0, 0.9, 0.72]
# α_t^DDPM = α_t / α_(t−1) and β_t = 1 − α_t / α_(t−1)
[round(alphas[t] / alphas[t - 1], 4) for t in (1, 2)] # → [0.9, 0.8]
[round(1 - alphas[t] / alphas[t - 1], 4) for t in (1, 2)] # → [0.1, 0.2]
D: experimental details
Architectures follow DDPM (a U-Net based on a Wide ResNet). LSUN results (Table 3) tell the same story as Table 1: on Bedroom, 20 DDIM steps give FID 8.89 against 22.77 for DDPM with η = 1, and with 100 steps they are close (6.62 and 6.81). Interpolations use slerp, and grids of interpolations blend four noises two ways.
Why it matters
The appendices are where the paper's generality lives: the discrete case hints that the same trick works beyond bell curves, and the notation map is what lets DDPM's trained checkpoints be reused without change.
What changed since 2020
The idea that training and sampling are separable has held up completely. What sits on either side of it has moved on:
| In the paper | Common today | Why | Where to read |
|---|---|---|---|
| DDIM with 20 to 100 steps | Higher-order ODE solvers, often 10 to 30 steps; distilled students with a handful | Every step is a network call | diffusion lesson |
| Unconditional or class samples | Text prompts with classifier-free guidance applied inside a DDIM-style step | Generate what was asked for | classifier-free guidance companion |
| Curved noising paths, stepped on a better clock | Flow matching: train the road to be straight | Straight roads forgive big steps | flow matching companion |
| Inversion for reconstruction | Inversion for editing real photos | Find a photo's noise, regenerate it differently | §5.4 above |
Glossary
Every term with hover guidance on this page, in one place.