An annotated companion · AI Primer

Denoising Diffusion Probabilistic Models, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The equations are reproduced because mathematics is not copyrightable. The paper is distributed under arXiv's standard licence, so its tables are not reproduced: a few selected rows appear in this page's own tables, each attributed, and its figures are redrawn from scratch. Worked numbers come from a two-step toy chain built on this page, not from the paper, unless a sentence says otherwise. Read the original alongside: every section links to it.

How to read this page

Nothing on this page assumes you already know the jargon. Three things help:

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same, and each equation is followed by a table of its symbols.
  • The diagrams are live: hover or tap any part to see what it does. The charts read out their values as you hover or use the arrow keys.

Each idea is explained in the same order: an everyday picture, a tiny example you could check by hand, a diagram, then the math, and finally why it still matters. Most of the math is checked on one tiny chain:

The two-step toy chain. One pixel with clean value x0 = 2.0. Step 1 mixes in a share β1 = 0.1 of noise, step 2 a share β2 = 0.2. The noise draws are ε1 = 0.5 and ε2 = −1.0, so the pixel goes 2.0 → 2.0555 → 1.3913. The paper uses the same arithmetic with 1,000 steps and whole images. The diffusion lesson builds all of it from scratch in NumPy, on dots in a plane instead of images.

Abstract

“We present high quality image synthesis results using diffusion probabilistic models, a class of latent variable models inspired by considerations from nonequilibrium thermodynamics.”Ho, Jain and Abbeel (2020), Abstract. Read the original

Everyday picture

A sculptor who cannot carve a horse in one go can still improve a rough horse a little. Chain that modest skill a thousand times, starting from a shapeless block, and a horse comes out. A diffusion model is that sculptor: it only ever learns to remove a little noise, and the practice material is free, because you can add noise to any real picture yourself and know the right answer.

What the paper claims

  • Diffusion models, proposed in 2015 but not yet shown to produce good images, can: on unconditional CIFAR10 (small 32 × 32 photos), an Inception score of 9.46 and a state-of-the-art FID of 3.17.
  • The best results come from a weighted version of the model's training bound, found through a new link to denoising score matching and Langevin dynamics.
  • The sampler can be read as a progressive, lossy decompression, a generalisation of autoregressive decoding.
  • On 256 × 256 LSUN images, sample quality similar to ProgressiveGAN.

Why it matters today

This paper's training loop, “noise a real image, ask a network for the noise, penalise the squared error”, is the loop inside Stable Diffusion's first versions and a long line of image, video and audio generators. Later work changed the sampler, the schedule and the network, but the idea started here.

1 Introduction · original

“A diffusion probabilistic model (which we will call a ‘diffusion model’ for brevity) is a parameterized Markov chain trained using variational inference to produce samples matching the data after finite time.”Ho, Jain and Abbeel (2020), §1

Everyday picture

Think of an old television losing its signal, one notch of static at a time, until only snow is left. That fixed, unlearned fading is the forward process. The model learns to run the same road backwards, from snow to a picture: the reverse process. Both are Markov chains: each step looks only at the step before.

Tiny example

In the toy chain, the forward process takes 2.0 to 2.0555 to 1.3913; each step used only the value just before it. The reverse process must go from 1.3913 back towards 2.0, one step at a time, without being told the noise.

The paper's three findings

  • Diffusion models can generate high-quality samples, sometimes better than other model families.
  • A particular way of setting up the network (predict the noise) is equivalent to denoising score matching over many noise levels when training, and to annealed Langevin dynamics when sampling. That parameterization gave the best samples.
  • Their log-likelihoods are not competitive: most of the bits go on details the eye cannot see, which the paper frames as lossy compression.

Why it matters

A generator trained by plain regression, with no opponent to balance as in a GAN, turned out to beat GANs on image quality. That stability is much of why diffusion took over; the GANs lesson shows the balancing act it avoided.

2 Background · original

Everyday picture

A diffusion model is a VAE stretched into a long chain. In a VAE, an encoder turns a picture into a fuzzy code and a decoder turns it back. Here the “encoder” is fixed arithmetic (add a little noise, a thousand times) and the “code” is pure noise of the same size as the picture. Only the decoder, the reverse chain, is learned. That makes the ELBO of the VAE paper the natural training objective.

xT ··· xt xt−1 ··· x0 learned step downp_θ(x_t−1 | x_t) fixed step up (dashed)q(x_t | x_t−1)

Hover or tap a circle or a label. Start with x0, the real picture, at the bottom.

Figure 2 of the paper, redrawn (turned upright to fit a phone): the diffusion chain. Based on Ho, Jain and Abbeel (2020), Figure 2.

Reading it: start at the bottom with x0, a real picture (shaded because it is observed). The dashed arrow is one forward step q, which adds a little noise and moves one place up; repeated T = 1,000 times, it reaches xT, pure noise, at the top. The solid arrows point the other way: each is one learned reverse step pθ, which removes a little noise and moves one place down. Generating a picture means starting from fresh noise at xT and following the solid arrows all the way down to x0. (The paper draws the same chain on its side, from xT on the left to x0 on the right.) Every xt has the same size as the picture: unlike a VAE's code, the “latents” here are never smaller than the data.

The reverse process: the model

Every reverse step is a bell curve whose centre and width a network computes from the current noisy picture and the step number:

In words: “the probability of a whole path from noise to picture is the probability of the starting noise times the probability of each reverse step; each step is a bell curve centred where the network says, with the spread the network says.”

With the numbers (illustrative reverse model: centres 1.90 and 1.95, variances β2 = 0.2 and β1 = 0.1): for the toy path 1.3913 → 2.0555 → 2.0, log p = log N(1.3913; 0, 1) + log N(2.0555; 1.90, 0.2) + log N(2.0; 1.95, 0.1) = −1.8867 − 0.1747 + 0.2199 = −1.8415. Logs of products are sums, which is why every bound below is a sum over steps.

In Python:

import math
def log_N(x, mean, var):
    # log of the bell-curve density
    return -0.5 * math.log(2 * math.pi * var) - (x - mean) ** 2 / (2 * var)
x0 = 2.0
# the toy path, from the forward process
x1 = math.sqrt(0.9) * x0 + math.sqrt(0.1) * 0.5
x2 = math.sqrt(0.8) * x1 + math.sqrt(0.2) * -1.0
round(x1, 4), round(x2, 4)  # → (2.0555, 1.3913)
# log p(x_T) + Σ_t log p_θ(x_t−1 | x_t), with illustrative centres 1.90 and 1.95
terms = [log_N(x2, 0, 1), log_N(x1, 1.90, 0.2), log_N(x0, 1.95, 0.1)]
[round(v, 4) for v in terms]  # → [-1.8867, -0.1747, 0.2199]
round(sum(terms), 4)  # → -1.8415

The forward process: fixed, no learning

In words: “each forward step shrinks the previous value a little, by √(1 − βt), and adds bell-curve noise of variance βt.”

With the numbers: step 1 of the toy chain has centre √0.9 × 2.0 = 1.8974 and variance 0.1; the draw ε1 = 0.5 lands at 1.8974 + √0.1 × 0.5 = 2.0555. Step 2: √0.8 × 2.0555 + √0.2 × (−1.0) = 1.8385 − 0.4472 = 1.3913.

In Python:

import math
x0, betas, eps = 2.0, [0.1, 0.2], [0.5, -1.0]
# the centre of q(x_1 | x_0): √(1 − β_1) x_0
round(math.sqrt(1 - betas[0]) * x0, 4)  # → 1.8974
x = x0
path = []
for beta, e in zip(betas, eps):
    # x_t = √(1 − β_t) x_t−1 + √β_t ε
    x = math.sqrt(1 - beta) * x + math.sqrt(beta) * e
    path.append(round(x, 4))
path  # → [2.0555, 1.3913]

The training bound

Exactly as in a VAE, the log-probability of the data cannot be computed, but a bound on it can. Here it is written as an upper bound on the negative log-likelihood, so training pushes it down:

In words: “run the forward chain from a real picture; score the noise at the end under the standard bell curve, and at every step compare how likely the reverse model finds going back with how likely the forward chain was to come forward. On average over many runs this never falls below the true negative log-likelihood.”

With the numbers (the toy path and the illustrative reverse model above): 1.8867 for the end noise, −0.4396 for step 2 and −0.1125 for step 1, a one-path estimate of L = 1.3347. Continuous log-densities can be positive, so terms can be negative; that is one reason §3.3 makes the last step discrete.

In Python:

import math
def log_N(x, mean, var):
    return -0.5 * math.log(2 * math.pi * var) - (x - mean) ** 2 / (2 * var)
x0 = 2.0
x1 = math.sqrt(0.9) * x0 + math.sqrt(0.1) * 0.5
x2 = math.sqrt(0.8) * x1 + math.sqrt(0.2) * -1.0
# −log p(x_T)
end = -log_N(x2, 0, 1)
# −log p_θ(x_1 | x_2) / q(x_2 | x_1), then −log p_θ(x_0 | x_1) / q(x_1 | x_0)
step2 = -(log_N(x1, 1.90, 0.2) - log_N(x2, math.sqrt(0.8) * x1, 0.2))
step1 = -(log_N(x0, 1.95, 0.1) - log_N(x1, math.sqrt(0.9) * x0, 0.1))
[round(v, 4) for v in (end, step2, step1)]  # → [1.8867, -0.4396, -0.1125]
round(end + step2 + step1, 4)  # → 1.3347

The shortcut: jump to any step

Everyday picture: adding noise twice is the same as adding it once in a bigger dose. So there is no need to run the chain step by step: multiply up how much signal survives, and jump.

In words: “after t steps, the noisy picture is a bell curve centred on the clean picture shrunk by the square root of the surviving signal share, with the rest of the variance filled by noise.”

With the numbers: the toy chain keeps 0.9 × 0.8 = 0.72, so q(x2 | x0) = N(√0.72 × 2.0, 0.28) = N(1.6971, 0.28). The paper's 1,000 steps (β rising in a straight line from 0.0001 to 0.02) keep ᾱ1000 = 0.00004: a signal share of 0.0064, practically nothing.

In Python:

import math
x0 = 2.0
alpha_bar = 1.0
# ᾱ_t = Π α_s, with α_s = 1 − β_s
for beta in [0.1, 0.2]:
    alpha_bar *= 1 - beta
round(alpha_bar, 2)  # → 0.72
# the centre and variance of q(x_2 | x_0)
round(math.sqrt(alpha_bar) * x0, 4), round(1 - alpha_bar, 2)  # → (1.6971, 0.28)
# the paper's schedule: 1,000 betas from 0.0001 to 0.02
T = 1000
betas = [1e-4 + (0.02 - 1e-4) * i / (T - 1) for i in range(T)]
alpha_bar_T = 1.0
for beta in betas:
    alpha_bar_T *= 1 - beta
round(alpha_bar_T, 5), round(math.sqrt(alpha_bar_T), 4)  # → (4e-05, 0.0064)

Reading it: the x-axis is the step t, from 1 to 1,000. The solid line is the signal share √ᾱt, how much of the clean picture is left; the dashed line is the noise share √(1 − ᾱt). With the paper's βT = 0.02 they cross near step 260 and the signal ends at 0.0064. Drag βT down and the signal survives longer, so the last step is no longer pure noise, and sampling would start from a picture that does not match the real end of the chain. Drag it up and the signal is gone halfway through, wasting steps. The readout below the chart updates as you hover. Computed live from the formula.

Splitting the bound into comparisons of bell curves

The paper rewrites L so that each term compares two bell curves directly, which can be done exactly rather than by sampling:

In words: “the bound is three kinds of cost: how far the end of the forward chain is from pure noise; at every middle step, how far the reverse model's step is from the true backward step (which is known exactly once you know the clean picture); and how well the last step rebuilds the picture.”

With the numbers: in the toy chain, LT = KL(N(1.6971, 0.28) ‖ N(0, 1)) = 1.7165: two steps leave far too much signal. The paper's schedule makes LT tiny, under 3 × 10−5 bits even for a pixel at the edge of its range (the paper reports ≈ 10−5 bits per dimension). The toy's middle term, with the illustrative reverse step N(1.90, 0.2), is L1 = 0.2185 (worked out under equation 8 below).

In Python:

import math
def kl_to_standard(mean, var):
    # D_KL(N(mean, var) ‖ N(0, 1))
    return 0.5 * (mean ** 2 + var - math.log(var) - 1)
# L_T for the toy chain: q(x_2 | x_0) = N(√0.72 · 2, 0.28)
round(kl_to_standard(math.sqrt(0.72) * 2.0, 0.28), 4)  # → 1.7165
# L_T for the paper's schedule, a pixel at x_0 = 1, in bits (divide nats by ln 2)
alpha_bar_T = 4.0358e-05
bits = kl_to_standard(math.sqrt(alpha_bar_T) * 1.0, 1 - alpha_bar_T) / math.log(2)
print(f"{bits:.1e}")  # → 2.9e-05

The true backward step, given the clean picture

Everyday picture: if you know where a walk started and where it is now, you can say quite precisely where it was one step ago. The backward step is hard only because the model does not know x0; during training we do, and the answer is a bell curve with a closed form:

In words: “one step back from xt, knowing x0, the value was a weighted blend of the clean picture and the current noisy one, give or take a spread slightly smaller than βt.”

With the numbers: toy chain, t = 2: the weights are √0.9 × 0.2 / 0.28 = 0.6776 on x0 and √0.8 × 0.1 / 0.28 = 0.3194 on x2, so μ̃2 = 0.6776 × 2.0 + 0.3194 × 1.3913 = 1.7997, and β̃2 = (0.1 / 0.28) × 0.2 = 0.0714. The true x1 was 2.0555, within about one spread (√0.0714 = 0.27) of that centre.

In Python:

import math
beta_t, alpha_t = 0.2, 0.8
alpha_bar_prev, alpha_bar_t = 0.9, 0.72
x0, xt = 2.0, 1.3913
# the two blending weights
c0 = math.sqrt(alpha_bar_prev) * beta_t / (1 - alpha_bar_t)
ct = math.sqrt(alpha_t) * (1 - alpha_bar_prev) / (1 - alpha_bar_t)
round(c0, 4), round(ct, 4)  # → (0.6776, 0.3194)
# μ̃_t and β̃_t
round(c0 * x0 + ct * xt, 4)  # → 1.7997
round((1 - alpha_bar_prev) / (1 - alpha_bar_t) * beta_t, 4)  # → 0.0714

Why it matters today

Every quantity on this section is closed-form, so training never simulates the chain. It picks a step, jumps there with the shortcut, and compares bell curves. The diffusion lesson draws the forward process and its schedule on dots in the plane.

In code: linear_schedule builds the list of βt as a NoiseSchedule, whose NoiseSchedule.alpha_bars is ᾱt; noise_step is one forward step and add_noise is the shortcut.

3 Diffusion models and denoising autoencoders · original

Everyday picture

The bound of §2 leaves many choices open: the schedule, the network, and how exactly the network describes each reverse step. This section makes those choices one term of the bound at a time: LT (§3.1), the middle terms L1:T−1 (§3.2), and L0 (§3.3). Along the way, the training loss turns into something a photo restorer would recognise: sprinkle known dust on a clean photo and grade the guess of where the dust went.

3.1 Forward process and LT · original

Everyday picture

If the static dial is bolted in place, how much static is on screen at the end is not something training can change. The paper fixes every βt to a constant, so the forward process has no learnable parameters.

Tiny example

The toy chain's LT = 1.7165 depends only on x0 and the fixed betas. No weight of the network appears in it, so its gradient is zero and training can drop it. (The paper notes the betas could be learned by reparameterization, as in the VAE paper, but chooses not to.)

Why it matters today

A fixed, parameter-free forward process is what separates diffusion models from VAEs: the “encoder” never has to be trained, and its end point is always the same known noise to sample from.

3.2 Reverse process and L1:T−1: guess the noise · original

“The complete sampling procedure, Algorithm 2, resembles Langevin dynamics with εθ as a learned gradient of the data density.”Ho, Jain and Abbeel (2020), §3.2

Everyday picture

First choice: do not learn the width of each reverse step. The paper fixes Σθ = σt²I and tries σt² = βt and σt² = β̃t, with similar results (the first is optimal if the data were pure noise, the second if the data were one single point). That leaves the centre. The obvious design asks the network for the centre directly. The paper's design asks it for something simpler: the noise that was added. A photo restorer who practises by guessing where known dust went is learning exactly this.

Tiny example

In the toy chain, x2 = 1.3913 is a one-jump mix of x0 = 2.0 and some total noise ε: 1.3913 = √0.72 × 2.0 + √0.28 × ε, so ε = −0.5779. Plug that ε into the paper's formula for the centre and you get 1.7997: exactly the true backward centre μ̃2 from §2. A network that guesses the noise perfectly gets the reverse step exactly right.

The middle terms are squared distances between centres

In words: “when both bell curves have fixed widths, the only part of each middle term training can change is the squared distance between the true backward centre and the model's centre, divided by twice the model's variance.”

With the numbers: toy chain, t = 2, σ2² = β2 = 0.2 and an illustrative model centre of 1.90: (1.7997 − 1.90)² / 0.4 = 0.0252. The full KL is 0.2185, and the gap, C = 0.1934, depends only on β̃2 and σ2², never on the network.

In Python:

import math
x0 = 2.0
x2 = math.sqrt(0.8) * (math.sqrt(0.9) * x0 + math.sqrt(0.1) * 0.5) + math.sqrt(0.2) * -1.0
# the true backward step from the previous block
mu_tilde = math.sqrt(0.9) * 0.2 / 0.28 * x0 + math.sqrt(0.8) * 0.1 / 0.28 * x2
beta_tilde = 0.1 / 0.28 * 0.2
mu_theta, sigma2 = 1.90, 0.2
# the part that depends on θ: ‖μ̃ − μ_θ‖² / (2σ²)
theta_part = (mu_tilde - mu_theta) ** 2 / (2 * sigma2)
round(theta_part, 4)  # → 0.0252
# the full KL between N(μ̃, β̃) and N(μ_θ, σ²)
kl = 0.5 * (beta_tilde / sigma2 + (mu_tilde - mu_theta) ** 2 / sigma2 - 1 - math.log(beta_tilde / sigma2))
round(kl, 4)  # → 0.2185
# C: what is left, the same whatever μ_θ is
round(kl - theta_part, 4)  # → 0.1934

Predict the noise instead of the centre

Write the noisy picture with the shortcut, xt = √ᾱt x0 + √(1 − ᾱt) ε, and substitute into μ̃t: the target centre becomes (xt − βt/√(1 − ᾱt) ε) / √αt. The network already sees xt, so the only unknown is ε. The paper's parameterization lets the network supply it:

In words: “to step back, subtract this step's share of the guessed noise, undo this step's shrink, then add a small fresh wobble.”

With the numbers: toy chain with the perfect guess ε = −0.5779: (1.3913 − 0.2/√0.28 × (−0.5779)) / √0.8 = (1.3913 + 0.2184) / 0.8944 = 1.7997 = μ̃2. The diffusion lesson's example step, xt = 1.0, βt = 0.19, ᾱt = 0.75 and a guess of 0.5, gives (1.0 − 0.38 × 0.5)/0.9 = 0.9, and with σt = √0.19 and z = 0.5 the next point is 1.1179.

In Python:

import math
x0 = 2.0
xt = math.sqrt(0.8) * (math.sqrt(0.9) * x0 + math.sqrt(0.1) * 0.5) + math.sqrt(0.2) * -1.0
beta_t, alpha_bar_t = 0.2, 0.72
alpha_t = 1 - beta_t
# the total noise that turned x_0 into x_t in one jump
eps = (xt - math.sqrt(alpha_bar_t) * x0) / math.sqrt(1 - alpha_bar_t)
round(eps, 4)  # → -0.5779
# μ_θ = (x_t − β_t / √(1 − ᾱ_t) · ε_θ) / √α_t, with a perfect guess ε_θ = ε
mu = (xt - beta_t / math.sqrt(1 - alpha_bar_t) * eps) / math.sqrt(alpha_t)
round(mu, 4)  # → 1.7997
# one sampling step from the diffusion lesson's example, σ_t = √β_t
x_t, b, ab, eps_hat, z = 1.0, 0.19, 0.75, 0.5, 0.5
round((x_t - b / math.sqrt(1 - ab) * eps_hat) / math.sqrt(1 - b) + math.sqrt(b) * z, 4)  # → 1.1179

Put this centre back into equation 8 and the middle terms become a weighted squared error on the noise:

In words: “noise a real picture to step t in one jump, ask the network what noise was added, and score the squared miss, weighted by a number that depends only on the step.”

With the numbers: with σt² = βt the weight simplifies to βt / (2αt(1 − ᾱt)). Toy chain, t = 2: 0.2 / (2 × 0.8 × 0.28) = 0.4464. Paper's schedule: 0.5 at t = 1, falling to 0.0102 at t = 1,000, a ratio of 49.

In Python:

beta_t, alpha_bar_t = 0.2, 0.72
sigma2 = beta_t
# β² / (2σ² α (1 − ᾱ)), with σ² = β
round(beta_t ** 2 / (2 * sigma2 * (1 - beta_t) * (1 - alpha_bar_t)), 4)  # → 0.4464
# the paper's schedule, first and last step
T = 1000
betas = [1e-4 + (0.02 - 1e-4) * i / (T - 1) for i in range(T)]
alpha_bars, ab = [], 1.0
for b in betas:
    ab *= 1 - b
    alpha_bars.append(ab)
w = [b / (2 * (1 - b) * (1 - a)) for b, a in zip(betas, alpha_bars)]
round(w[0], 3), round(w[-1], 4), round(w[0] / w[-1])  # → (0.5, 0.0102, 49)

This is the paper's central observation. The loss “resembles denoising score matching over multiple noise scales”: guessing the noise is guessing the score, the direction towards more data, flipped and rescaled. And the sampling step, “move a little along a learned direction, then add a little noise”, is the shape of Langevin dynamics. Training this bound trains that sampler.

Algorithm 1: training Algorithm 2: sampling repeat until converged x₀ t ε x_t = √ᾱ x₀ + √(1−ᾱ) ε network ε_θ(x_t, t) ‖ε − ε_θ‖² gradient step t ← t − 1 x_T ~ N(0, I) guess ε_θ(x_t, t) compute μ_θ + σ_t z (if t > 1) x₀

Hover or tap a step. Training is on the left, sampling on the right.

Algorithms 1 and 2 of the paper, redrawn as flow diagrams. Based on Ho, Jain and Abbeel (2020), Algorithms 1 and 2.

Reading it: on the left, training takes three random choices at the top (a real picture, a step, a noise draw), mixes them with the shortcut into one noisy picture, and asks the network for the noise. Only the noise draw also takes the long wire down the right-hand side, straight to the loss, where the guess is graded against it. The whole column repeats until the loss stops improving; there is no chain to simulate. On the right, sampling starts from pure noise and loops T times: guess the noise, compute the centre μθ, add a fresh wobble σtz (except on the last step), step t down by one. Each lap is one full run of the network, so a sample costs T = 1,000 network calls.

Why it matters today

“Predict the noise” is still the most common parameterization, and the two algorithms are the ones a first diffusion implementation writes. Later samplers such as DDIM reuse the same trained εθ with far fewer steps; see the diffusion lesson.

In code: train_noise_predictor is Algorithm 1 on a small DenoiserMLP; ddpm_step is one step of Algorithm 2 with σt = √βt, and ddpm_sample runs the loop; noise_to_score turns the noise guess into the score.

3.3 Data scaling, reverse process decoder, and L0 · original

Everyday picture

Real pixels are whole numbers from 0 to 255, not smooth values. So the very last step does not ask “how likely is exactly this value?” (a question with no good answer for a continuous bell curve) but “how much of the bell curve falls in this pixel's bin?”, the way a histogram counts everything inside a bar.

Tiny example

Pixel values are scaled linearly from {0, …, 255} to [−1, 1]. A pixel of 128 becomes 128/127.5 − 1 = 0.0039, and its bin runs half a step either side, from 0.0000 to 0.0078. If the model's last step is a bell curve centred at 0.01 (illustrative) with spread σ1 = √β1 = 0.01 (the paper's schedule), the chance of landing in that bin is 0.256.

In words: “the probability of the final picture is, for every pixel independently, the area under that pixel's bell curve over the pixel's bin; the outermost bins stretch to infinity so no probability is lost off the ends.”

With the numbers: Φ((0.0078 − 0.01)/0.01) − Φ((0.0000 − 0.01)/0.01) = Φ(−0.22) − Φ(−1.0) = 0.256 for the pixel of 128. Multiply such areas over all D = 32 × 32 × 3 = 3,072 numbers of a CIFAR10 image to get pθ(x0 | x1).

In Python:

from statistics import NormalDist
# scale a pixel from {0, ..., 255} to [−1, 1]
x = 128 / 127.5 - 1
round(x, 4)  # → 0.0039
# the bin: δ−(x) = x − 1/255, δ+(x) = x + 1/255
lo, hi = x - 1 / 255, x + 1 / 255
round(lo, 4), round(hi, 4)  # → (-0.0, 0.0078)
# area of N(μ, σ_1²) over the bin, with an illustrative μ = 0.01 and σ_1 = √0.0001 = 0.01
step = NormalDist(0.01, 0.01)
round(step.cdf(hi) - step.cdf(lo), 3)  # → 0.256
# D for a CIFAR10 image
32 * 32 * 3  # → 3072

Why it matters today

This makes the bound an honest codelength in bits per dimension, comparable with other likelihood models. At the end of sampling, the paper shows μθ(x1, 1) without adding the last bit of noise.

3.4 The simplified training objective · original

Everyday picture

The exact bound grades every step's homework with a different weight, and it weights the easiest homework (removing a whisper of noise from a nearly clean picture) the most. The paper's simplification: grade every step equally. Students then spend their effort on the hard exercises, the heavily noised pictures, where the big decisions about the image are made.

Tiny example

True noise (0.5, −1.0), network guess (0.4, −0.8): the miss is (0.1, −0.2) and the squared miss 0.01 + 0.04 = 0.05. Under the exact bound this miss would be multiplied by the step's weight (0.5 at step 1, 0.0102 at step 1,000). Under Lsimple it counts as 0.05 at every step.

In words: “pick a random step, a random real picture and random noise; noise the picture to that step in one jump; score the squared distance between the real noise and the network's guess; average over everything.”

With the numbers: ‖(0.5, −1.0) − (0.4, −0.8)‖² = 0.05. A network that always guesses zero scores the average of ε1² + ε2², which is 2 for 2-number noise: the score to beat.

In Python:

import random
eps, guess = [0.5, -1.0], [0.4, -0.8]
# ‖ε − ε_θ‖²
round(sum((e - g) ** 2 for e, g in zip(eps, guess)), 2)  # → 0.05
# always guessing 0: the average of ‖ε‖² over many draws of 2-number noise
random.seed(0)
draws = [random.gauss(0, 1) ** 2 + random.gauss(0, 1) ** 2 for _ in range(100_000)]
round(sum(draws) / len(draws), 1)  # → 2.0

Hover or use the arrow keys to read the weight at any step.

Reading it: the x-axis is the step t on a logarithmic scale, so the first few steps get room; the y-axis is the weight the exact bound puts on the noise-guessing error at that step (with σt² = βt, the paper's schedule). It starts at 0.5 at t = 1, halves almost at once, and settles near 0.01 for most of the chain. Lsimple replaces this whole curve with a flat line: every step counts the same. Relative to the exact bound, that shifts attention from the first handful of steps (almost clean pictures) to the hundreds of noisy ones, which is the down-weighting of small t the paper describes. Computed live from the schedule.

Why it matters today

Lsimple gave the paper's best samples (FID 3.17 against 13.51 for the exact bound; §4.2), and plain, unweighted noise regression became the default. Later work returned to the question of how to weight the steps, but the unweighted loss is still the baseline every paper compares against.

In code: train_noise_predictor trains on exactly this loss, averaged over a batch of random steps, pictures and noises.

4 Experiments · original

Everyday picture

A recipe is only as good as the dish. The experiments cook the recipe on small photos (CIFAR10), faces (CelebA-HQ) and scenes (LSUN), then judge the results by standard image-quality scores and by the bits the model needs to describe a picture.

The setup

Settings stated in §4 and Appendix B of Ho, Jain and Abbeel (2020), collected into one table
SettingValue
Steps T1,000 (to match earlier work, without a sweep)
Scheduleβ rising linearly from 10−4 to 0.02, chosen from constant, linear and quadratic schedules
Networka U-Net like PixelCNN++'s, with group normalization, self-attention at the 16 × 16 resolution, and the step given by the Transformer's sinusoidal embedding; one network shared by all steps
Size35.7 million parameters for CIFAR10; 114 million for LSUN and CelebA-HQ; about 256 million for a larger LSUN Bedroom model
OptimiserAdam, learning rate 2 × 10−4 (2 × 10−5 at 256 × 256); batch 128 for CIFAR10 and 64 for larger images; an exponential moving average of the weights with decay 0.9999
Regularisationdropout 0.1 on CIFAR10, random horizontal flips
Coston a TPU v3-8, CIFAR10 trains at 21 steps per second, 10.6 hours for 800,000 steps; sampling 256 images takes 17 seconds

Why it matters

Nothing exotic: a standard image network, a standard optimiser, a simple schedule. The quality came from the objective and the parameterization. See the Adam companion and, for the attention blocks and sinusoidal step embedding, the Transformer companion.

4.1 Sample quality · original

Everyday picture

How do you grade an artist who paints pictures nobody asked for? FID compares the whole gallery of generated pictures with a gallery of real ones, as clouds of features from a pretrained image network: the closer the clouds, the lower the score. Lower is better; 0 would be indistinguishable.

Tiny example

Two galleries whose feature clouds have the same shape but whose centres sit 3 units apart score an FID of 3² = 9; move the centres to 1 unit apart and it drops to 1. (FID also compares the clouds' spreads; with equal spreads that part is zero.)

Selected rows of Table 1, CIFAR10, reproduced with attribution (Ho, Jain and Abbeel, 2020). NLL in bits per dimension, test (train)
ModelInception scoreFIDNLL
BigGAN (conditional)9.2214.73
Gated PixelCNN4.6065.933.03 (2.90)
Sparse Transformer2.80
NCSN (score matching)8.8725.32
StyleGAN2 + ADA (v1)9.743.26
Diffusion, exact bound L7.6713.51≤ 3.70 (3.69)
Diffusion, Lsimple9.463.17≤ 3.75 (3.72)

Reading it: each bar is one model's FID on unconditional CIFAR10 (BigGAN is class-conditional, so it had extra help), drawn to one scale from 0 to 70; shorter is better. The striped bars are the paper's two diffusion models. Trained on the exact bound, the model sits among earlier methods at 13.51. Trained on Lsimple, it reaches 3.17, just ahead of StyleGAN2 + ADA's 3.26. Its FID against the test set rather than the training set is 5.24. The NLL column shows the other side: the Sparse Transformer's 2.80 bits per dimension beats both diffusion models' 3.70 and 3.75.

Why it matters today

This table is why diffusion models took off: a likelihood-based model, trained by regression, matching the best GANs on the standard image benchmark. On 256 × 256 LSUN, the paper reports FID 7.89 on churches and 4.90 on bedrooms (larger model).

4.2 Reverse process parameterization and training objective ablation · original

Everyday picture

Two knobs, four settings, one winner. Knob one: does the network predict the centre μ̃ or the noise ε? Knob two: train on the exact bound, or on unweighted squared error? And a side question: should the network also learn the width of each step?

Selected rows of Table 2, unconditional CIFAR10, reproduced with attribution (Ho, Jain and Abbeel, 2020)
PredictObjectiveInception scoreFID
centre μ̃exact bound, learned widths7.2823.69
centre μ̃exact bound, fixed widths8.0613.22
noise εexact bound, fixed widths7.6713.51
noise εLsimple9.463.17

What the rows say

  • Predicting the centre works only on the exact bound; trained with unweighted squared error on the centre it was unstable.
  • Learning the widths made training unstable and samples worse, for both parameterizations.
  • On the exact bound, predicting the noise is about as good as predicting the centre (13.51 against 13.22). With Lsimple, it is far better (3.17).

Why it matters today

The winning combination is a pair: ε-prediction and the unweighted loss. The paper also mentions trying to predict x0 directly, which gave worse samples early on. Later work revisited learned widths with a different parameterization, but the ε-prediction, fixed-width recipe of this table is where most implementations start.

4.3 Progressive coding · original

“More than half of the lossless codelength describes imperceptible distortions.”Ho, Jain and Abbeel (2020), §4.3

Everyday picture

Picture a slow internet connection loading a photo: first a blurry thumbnail, then sharper and sharper. The bound can be read as exactly such a transmission: send xT, then xT−1, and so on, each costing the bits of one KL term. At any moment the receiver can guess the final picture from what has arrived.

Tiny example

Midway, the receiver holds the toy's xt = 1.9 at a level with ᾱt = 0.64, and the network guesses the noise in it as 0.5. Undoing the shortcut gives the receiver's best guess of the clean value: (1.9 − 0.6 × 0.5) / 0.8 = 2.0.

In words: “the best guess of the clean picture is the noisy one with the guessed noise taken out, scaled back up by the surviving signal share.”

With the numbers: (1.9 − √0.36 × 0.5) / √0.64 = (1.9 − 0.3) / 0.8 = 2.0.

In Python:

import math
x_t, alpha_bar_t, eps_hat = 1.9, 0.64, 0.5
# x̂_0 = (x_t − √(1 − ᾱ_t) ε_θ) / √ᾱ_t
round((x_t - math.sqrt(1 - alpha_bar_t) * eps_hat) / math.sqrt(alpha_bar_t), 4)  # → 2.0
CIFAR10 test set, from the values in the paper's Table 4.

Hover or use the arrow keys to read a point.

Hover or use the arrow keys to read a point.

Reading it: both charts share an x-axis, the reverse-process time T − t + 1 as the paper's Table 4 lists it, from step 100 to step 1,000 of generation. The top chart is distortion, the root-mean-squared error of the receiver's guess x̂0 on a 0 to 255 pixel scale: it falls steadily from 67.6 to 12.0 by step 900 and to 0.95 at the end. The bottom chart is rate, the bits per dimension received so far: almost nothing until step 800 (0.054), 0.12 at step 900, and then a jump to 1.78 in the last hundred steps. So most of the bits buy the last drop in error, from 12 grey levels to under 1: details the eye barely sees. This is Figure 5 of the paper, redrawn from its Table 4.

The paper's best CIFAR10 model has a rate of 1.78 bits per dimension and a distortion of 1.97 bits per dimension, a root-mean-squared error of 0.95 on the 0 to 255 scale. Its train and test codelengths differ by at most 0.03 bits per dimension, so it is not overfitting.

Progressive generation

Running the sampler and showing x̂0 along the way, large features (layout, colour) appear first and fine details last. Freezing xt and sampling several endings shows the same thing from the other side: when t is large only the coarse features are shared; when t is small, all but the finest details are.

A connection to autoregressive models

Tiny example: take a 3-pixel image and a strange “diffusion” that blanks one pixel per step: step 1 blanks pixel 1, step 2 pixel 2, step 3 pixel 3, ending at a blank image. Running it backwards means predicting pixel 3, then pixel 2 given pixel 3, then pixel 1 given the other two: that is an autoregressive model. The paper shows the bound becomes exactly an autoregressive loss under that choice, and so reads Gaussian diffusion as autoregression along a generalised “bit ordering” that no reordering of pixels could express. T also need not equal the number of pixels: T = 1,000 is less than the 3,072 numbers of a CIFAR10 image.

Why it matters today

The coarse-to-fine behaviour is why step budgets matter: cut the early steps and the layout suffers, cut the late ones and only fine texture does. It also explains why diffusion models make better pictures than their likelihoods suggest: the likelihood pays for invisible details. The paper notes Algorithms 3 and 4 are a proof of concept, since the coding procedure they assume is not tractable for high-dimensional data.

4.4 Interpolation · original

Everyday picture

Blend two faces by mixing their pixels and you get a ghostly double exposure. Instead, noise both faces part of the way, blend the noisy versions, and let the reverse process clean up the blend: it removes the ghosting and paints a single, plausible face between the two.

Tiny example

Two pictures noised to the same step (with the same noise, so only the pictures differ) have values 1.9 and −0.5 at one pixel. A quarter of the way from the first to the second: 0.75 × 1.9 + 0.25 × (−0.5) = 1.3. The reverse process then denoises that blended value.

In words: “take a share 1 − λ of the first noised picture and a share λ of the second, then run the reverse process from the blend.”

With the numbers: λ = 0.25: 0.75 × 1.9 + 0.25 × (−0.5) = 1.3. The printed formula in the paper writes the clean pictures x0 and x′0 on the right; the surrounding text (“the linearly interpolated latent”, with the noise fixed so that xt and x′t stay the same) makes clear the noised versions are meant, as written here.

In Python:

x_t, x_t_prime, lam = 1.9, -0.5, 0.25
# x̄_t = (1 − λ) x_t + λ x′_t
round((1 - lam) * x_t + lam * x_t_prime, 4)  # → 1.3

What the paper found

On CelebA-HQ faces at t = 500, interpolations vary pose, skin tone, hairstyle, expression and background smoothly, though not eyewear. More noising steps before blending give coarser, more varied interpolations; after 1,000 steps the source pictures are forgotten and the results are new samples.

Why it matters today

“Noise part of the way, then denoise” became a standard editing tool: image-to-image translation and sketch-to-picture tools use exactly this move, choosing how far to noise as a strength knob.

5 Related work · original

Everyday picture

Diffusion models look like cousins of several families at once: of VAEs (a latent-variable model trained on a bound), of flows (a chain of transformations), and of score-based models (following learned arrows towards the data). The paper places itself among them.

What it says

  • Versus VAEs and flows: the “encoder” q has no parameters, and the top latent xT shares almost no information with the data. See the VAE companion for the bound both families train on.
  • Versus score matching (NCSN): the same network idea, but diffusion admits a straightforward log-likelihood and trains its sampler directly by variational inference. Appendix C lists four differences: a U-Net with self-attention; shrinking the data by √(1 − βt) each step so the network's inputs stay the same size; a forward process that truly destroys the signal (LT ≈ 0) with small βt; and sampler coefficients derived from βt rather than set by hand.
  • The reverse implication: a certain weighted form of denoising score matching is variational inference for a Langevin-like sampler.

Tiny example

The shrink matters. Without it, adding noise of variance 0.1 and then 0.2 to a pixel of variance 1 leaves variance 1.3, and after many steps the inputs grow without limit. With it, (1 − 0.1) × 1 + 0.1 = 1 and then (1 − 0.2) × 1 + 0.2 = 1: the variance stays at 1 at every step.

Why it matters today

The two views, “chain of denoising steps” and “follow the score”, were soon unified as one family described in continuous time. The diffusion lesson covers the score view and its later relative, flow matching.

6 Conclusion and broader impact · original

What the paper concludes

Diffusion models produce high-quality images, and they connect to variational inference for Markov chains, denoising score matching, annealed Langevin dynamics, energy-based models, autoregressive models and progressive lossy compression. The authors expect good inductive biases for image data and look to other data types and to diffusion as a component in other systems.

Broader impact

The paper names the risks directly: generative models make fake images and videos of public figures easier to produce, and they reflect and can reinforce the biases of datasets scraped from the internet. It also names benefits: compression, representation learning for downstream tasks, and creative uses in art, photography and music.

Why it matters today

Both halves came true within a few years. Text-to-image, video and audio generators are built on this recipe, and detecting and labelling generated media became a field of its own.

What changed since 2020

The training loop of Algorithm 1 is recognisably the same. Almost everything around it has been tuned:

Choice in the paperCommon todayWhyWhere to read
1,000 sampling steps (Algorithm 2)DDIM and other samplers, 20 to 50 steps; distilled models, a handfulA sample costs one network call per stepdiffusion
Diffusion on pixelsLatent diffusion inside a VAE's codeTens of times fewer numbers per imageVAE companion
Unconditional samplesText conditioning with classifier-free guidanceGenerate what was asked fordiffusion
U-Net denoiserOften a transformer over latent patchesScales better with sizeTransformer companion
Curved noising paths, ε-predictionAlso flow matching with straight pathsFewer steps for the same qualitydiffusion

Glossary

Every term with hover guidance on this page, in one place.