An annotated companion · AI Primer

High-Resolution Image Synthesis with Latent Diffusion Models, annotated

About this page. This is a companion, not a copy. It follows the paper (version 2, April 2022) section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The equations are reproduced because mathematics is not copyrightable. The paper is distributed under arXiv's standard licence, so its tables are not reproduced: a few selected rows appear in this page's own tables and charts, each attributed, and its figures are redrawn from scratch. Worked numbers come from small examples built on this page, not from the paper, unless a sentence says otherwise; made-up numbers are labelled illustrative. Read the original alongside: every section links to it.

How to read this page

Nothing on this page assumes you already know the jargon. Three things help:

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same, and each equation is followed by a table of its symbols.
  • The diagrams are live: hover or tap any part to see what it does. The charts read out their values as you hover or use the arrow keys.

Each idea is explained in the same order: an everyday picture, a tiny example you could check by hand, a diagram, then the math, and finally why it still matters. This paper builds on two others with companions of their own: the DDPM companion (how a diffusion model is trained to guess noise) and the VAE companion (how an autoencoder learns a small, well-behaved code). Most of the arithmetic below uses one running example:

The running example. A colour photo of 256 × 256 pixels is 256 × 256 × 3 = 196,608 numbers. The paper's text-to-image autoencoder shrinks each side by f = 8 and keeps c = 4 numbers per position: a latent of 32 × 32 × 4 = 4,096 numbers, 48 times fewer. The diffusion model then does all of its work on that small grid. The diffusion lesson builds the diffusion half from scratch in NumPy, and the autoencoders lesson builds the autoencoder half.

Abstract

“To enable DM training on limited computational resources while retaining their quality and flexibility, we apply them in the latent space of powerful pretrained autoencoders.”Rombach et al. (2021), Abstract. Read the original

Everyday picture

An architect does not design a skyscraper by placing every brick. They draw a floor plan, small enough to think about, and a builder turns the plan into a building. A latent diffusion model splits image making the same way: an autoencoder learns once how to turn pictures into small “plans” and back, and a diffusion model learns to invent new plans. The expensive creative loop runs on the plan, never on the pixels.

What the paper claims

  • Diffusion in a mildly compressed latent space reaches “a near-optimal point between complexity reduction and detail preservation”: far cheaper to train and sample, with almost no loss of quality.
  • Cross-attention layers turn the denoiser into a general conditional generator: text, bounding boxes, class labels.
  • New state-of-the-art scores for inpainting and class-conditional ImageNet, and competitive text-to-image, unconditional generation and super-resolution, at a fraction of the compute of pixel-space diffusion.

Why it matters today

The released model family became Stable Diffusion. “Compress with an autoencoder, diffuse in the latent, decode once” is still the default shape of image and video generators, even where the denoiser has since become a transformer (see the DiT companion).

1 Introduction · original

“Most bits of a digital image correspond to imperceptible details.”Rombach et al. (2021), Figure 2 caption

Everyday picture

Imagine describing a photo over the phone. The first few sentences carry almost everything that matters: “a woman in a blue hat, smiling, outdoors”. After that you are reading out the exact texture of each thread in the hat, which takes far longer and which the listener would never miss. A diffusion model trained on pixels spends most of its effort on the thread-by-thread part, because every pixel is in its input and output, even though its loss already cares little about those details.

Tiny example: what pixel-space diffusion costs

The paper quotes the price of the strongest pixel-space models of the time: 150 to 1,000 V100 days to train, and about 5 days on one A100 card to draw 50,000 samples. That is about 8.6 seconds per image, and each image runs the same large network 25 to 1,000 times in a row.

In Python:

# 50,000 samples in about 5 days on one GPU
seconds = 5 * 24 * 60 * 60
round(seconds / 50_000, 1)  # → 8.6

Two stages of learning, drawn

The paper's argument starts from a trained pixel-space diffusion model (the DDPM one) read as a compressor. As its reverse process runs, you can count how many bits it has “sent” (the rate) and how far its current best guess of the final picture is from the truth (the distortion).

Hover or use the arrow keys to read a point.

Reading it: the x-axis is the rate, in bits per dimension received so far; the y-axis is the distortion, the root-mean-squared error of the current guess on a 0 to 255 pixel scale. Start at the top left: the first 0.12 bits per dimension (reverse steps 100 to 900) bring the error down from 67.6 to 12.0 grey levels. That steep drop is what the paper calls semantic compression: layout, shapes, what the picture is of. Then the curve lies almost flat: the last 1.66 bits per dimension, more than nine tenths of the total, only take the error from 12.0 to 0.95. That long tail is perceptual compression: detail the eye barely registers. The paper's point is that a separate, cheaper model (an autoencoder) can handle the tail, so the diffusion model need only learn the steep part. Its Figure 2 draws the same curve from the same model; the ten points here are the values the DDPM paper publishes in its Table 4.

The paper's plan and contributions

Train in two separate phases. First, an autoencoder that is “perceptually equivalent” to the image: its reconstructions look the same, but it holds far fewer numbers. Second, a diffusion model in that latent space. The authors list six contributions: milder compression than earlier two-stage methods, so reconstructions stay faithful; competitive results on many tasks at lower cost; no delicate balancing of reconstruction against generation, because the two are trained separately; convolutional sampling of images around 1024² pixels; a general cross-attention conditioning mechanism; and released models.

Why it matters

This separation is what put high-resolution diffusion on a single GPU. The autoencoder is trained once and reused for every diffusion model trained on top of it, which is also why later models could swap the denoiser while keeping the same latent space (the DiT paper reuses this paper's f = 8 autoencoder, with only its decoder fine-tuned).

2 Related work · original

Everyday picture

Before this paper, every family of image generator had one weak spot. GANs draw sharp pictures in one shot but are hard to train and can drop whole kinds of image (mode collapse). VAEs and flows train steadily but look blurrier. Autoregressive transformers are strong but read an image one token at a time, which limits them to small images. Diffusion models are strong and stable but slow and costly on pixels.

Two-stage models, and why they compressed so hard

Earlier two-stage systems (VQ-VAEs, VQGAN, DALL-E) also compressed images first, into a grid of codebook entries, then modelled that grid with an autoregressive transformer. A transformer's self-attention compares every token with every other, so its cost grows with the square of the number of tokens, and those systems had to squeeze hard (f = 16 or more) to keep the sequence short. Harsh squeezing throws away detail the decoder cannot recover.

Tiny example: a 256 × 256 image at f = 16 is a 16 × 16 grid, 256 tokens and 65,536 attention pairs per layer. At f = 8 it is 1,024 tokens and 1,048,576 pairs: 16 times the attention work for halving f. The convolutional U-Net this paper uses does work proportional to the number of positions, not its square, so it can afford the milder f.

In Python:

side = 256
for f in (16, 8):
    # tokens in an f-times-smaller grid, and attention pairs (every token with every token)
    tokens = (side // f) ** 2
    print(f, tokens, tokens ** 2)  # → 16 256 65536 8 1024 1048576
# halving f: how much more attention work
1048576 // 65536  # → 16

Why it matters

The paper's position: “LDMs scale more gently to higher dimensional latent spaces due to their convolutional backbone”, so they can choose the compression that best balances a faithful autoencoder against the work left to the generator. A jointly trained alternative (LSGM) exists, but needs a delicate weighting of reconstruction against generation; this paper keeps the two apart. The ViT companion covers the square-law cost of attention over image patches, which returns in the DiT companion.

3 Method · original

“We propose to circumvent this drawback by introducing an explicit separation of the compressive from the generative learning phase.”Rombach et al. (2021), §3

Everyday picture

A diffusion model can already ignore invisible detail through its loss, but it still has to look at every pixel on every step, both while training and while sampling. That is like a sculptor who has learned not to care about the grain of the marble, yet still has to walk around every square millimetre of it before each chip. Hand them a small clay maquette instead, and let a craftsman scale it up at the end.

Three advantages the paper lists

  • Cheaper diffusion: sampling happens on the small latent, so each of the many network calls is cheaper.
  • The right assumptions: the latent is still a 2-D grid, so the denoiser can be a convolutional U-Net with its built-in inductive bias for images, which is why mild compression suffices.
  • A reusable space: the same autoencoder serves many generative models and other uses.
Pixel space Latent space Pixel space Conditioning diffusion process (fixed) ×(T − 1) when the loop ends K, V x ℰ encoder z z_T denoising U-Net ε_θ QK V QK V QK V QK V step z_(T−1) 𝒟 decoder x̃ text, map, image τ_θ switch: concat

Hover or tap any part. Start at x, the photo, at the top left, and follow the arrows.

Figure 3 of the paper, redrawn (turned upright to fit a phone): conditioning a latent diffusion model by concatenation or by cross-attention. Based on Rombach et al. (2021), Figure 3.

Reading it: start at the top left. A photo x enters the encoder ℰ and becomes the small latent z, the first box in the big Latent space frame. During training, the fixed diffusion process (the long arrow) noises z towards pure noise zT. Generation runs the other way: zT enters the denoising U-Net from the right, the U-Net removes a little noise, and its output zT−1 loops back in, T − 1 times in all. The four purple blocks inside the U-Net are its attention layers: their queries Q come from the image features, while their keys and values K, V arrive from below, from the conditioning encoder τθ, which has turned a prompt (text, a layout, a class) into vectors. The dashed wire is the other way to condition: for inputs lined up with the image (a low-resolution photo, a semantic map), the switch simply stacks them onto zT as extra channels. When the loop ends, the clean latent leaves the frame at the bottom left and the decoder 𝒟 turns it into the picture x̃, once. Only the part inside the latent frame runs many times.

3.1 Perceptual image compression · original

Everyday picture

A good travel sketch keeps what matters about a street and drops the brickwork. The autoencoder is trained to make sketches that a skilled painter (the decoder) can turn back into something that looks like the photo, even if not every brick is where it was. “Looks like” is judged the way people judge: by a perceptual loss and by a small discriminator that inspects patches for realism, not by raw pixel-by-pixel error, which rewards blur.

Tiny example

The running example: x is 256 × 256 × 3. The encoder downsamples each side by f = 8, so h = w = 32, and keeps c = 4 channels. The latent z holds 4,096 numbers against the image's 196,608, a factor of 48. Shrink harder, f = 16 with c = 8, and it is 2,048 numbers; milder, f = 4 with c = 3, and it is 12,288.

In words: “the encoder turns an H × W colour image into an h × w grid with c numbers per cell, f times smaller on each side, where f is a power of two; the decoder turns such a grid back into an image.”

With the numbers: H = W = 256 and f = 23 = 8 give h = w = 32; with c = 4, z is 32 × 32 × 4 = 4,096 numbers, and 196,608 / 4,096 = 48.

In Python:

H, W, m, c = 256, 256, 3, 4
# f = 2^m, and h = H / f, w = W / f
f = 2 ** m
h, w = H // f, W // f
h, w  # → (32, 32)
# numbers in the image and in the latent
H * W * 3, h * w * c  # → (196608, 4096)
(H * W * 3) / (h * w * c)  # → 48.0

Two ways to keep the latent tame. Left alone, an autoencoder may spread its latent over huge values, which a diffusion model then handles badly. The paper tries two gentle regularizers. KL-reg. adds a tiny KL penalty pulling the latent towards the standard normal, as in a VAE, but weighted by only about 10−6. VQ-reg. puts a vector quantization layer inside the decoder, with a large codebook. Because the diffusion model sees the latent as a 2-D grid rather than a 1-D sequence, both can stay mild.

The full autoencoder objective is in the paper's Appendix G. The encoder and decoder minimize it; the patch discriminator Dψ maximizes it:

In words: “the autoencoder wants its rebuild to match the photo, to look real to the patch critic, and to keep its latent near a tame shape; the critic wants to call real photos real and rebuilds fake.”

With the numbers (illustrative): a rebuild error of 0.05, an adversarial term of 0.9, a critic that gives the real photo probability 0.8 (log 0.8 = −0.223), and a KL of 40 weighted by 10−6: 0.05 − 0.9 − 0.223 + 0.00004 = −1.073. The regularizer's share is 0.00004: it only stops the latent drifting to extreme scales, and barely competes with making good rebuilds.

In Python:

import math
# illustrative values for one image
L_rec, L_adv, D_real = 0.05, 0.9, 0.8
# L_reg: a KL of 40, weighted by 1e-6
L_reg = 1e-6 * 40
round(L_rec - L_adv + math.log(D_real) + L_reg, 3)  # → -1.073
round(L_reg, 6)  # → 4e-05

Hover or use the arrow keys to read a point.

Reading it: the slider picks how much the autoencoder shrinks each side. The square on the left is a 256 × 256 image; the grid on the right is the latent the denoiser will see, one cell per latent position (drawn as a shaded square when the cells are too small to draw). The readout counts the numbers per image and the shrink factor, with the reconstruction quality of the paper's KL-regularized autoencoder at that setting. The chart below plots that quality: PSNR (higher means a closer pixel match) and reconstruction FID (lower means the rebuilds look as natural as real images). Each step up in f roughly halves the numbers and costs about 2 to 5 dB of PSNR; reconstruction FID stays below 1 up to f = 8, reaches 2.63 at f = 16 and 7.3 at f = 32. At f = 1 there is no autoencoder: diffusion on pixels, perfect “reconstruction” and no saving. Values are selected rows of the paper's Table 8 (KL-regularized models, evaluated on ImageNet), reproduced with attribution.

Selected rows of Table 8 (and Figure 1) of Rombach et al. (2021): three autoencoders, reconstruction quality on ImageNet validation images
AutoencoderfR-FID ↓PSNR ↑
this paper, VQ-reg., |𝒵| = 8192, c = 340.5827.43
DALL-E's autoencoder832.0122.8
VQGAN, |𝒵| = 16384164.9819.9

Why it matters today

The autoencoder is where most of an image generator's fine detail comes from, and its latent shape (f = 8, c = 4 in this paper's text model) became a de facto standard that later models reused or widened. Everything the diffusion model draws must pass through this decoder, so its limits are the whole system's limits (§5). The autoencoders lesson builds a VAE and its KL term from scratch.

In code: latent_shrink counts pixels against latent numbers; VAE is a small KL-regularized autoencoder, and kl_to_standard_normal is its KL term.

3.2 Latent diffusion models · original

Everyday picture

Nothing about the diffusion recipe changes. The photo restorer still practises by sprinkling known dust on clean pictures and guessing where it went. The only difference is that the “pictures” are now the autoencoder's small sketches, so every practice round is 48 times cheaper.

Tiny example

Take one latent number whose clean value is z = 2.0, at a noise level where the signal share is √ᾱt = 0.8 and the noise share √(1 − ᾱt) = 0.6. With noise ε = 0.5, the noisy latent is zt = 0.8 × 2.0 + 0.6 × 0.5 = 1.9. The network sees 1.9 and the step t and guesses the noise. For two latent numbers with true noise (0.5, −1.0) and a guess of (0.4, −0.8), the squared miss is 0.01 + 0.04 = 0.05.

The paper first states the pixel-space objective of DDPM, then the same objective in the latent:

In words: “encode a real photo, noise its latent to a random step in one jump, ask the network what noise was added, and score the squared miss; average over photos, noise and steps. The first line is the same thing on pixels.”

With the numbers: zt = 0.8 × 2.0 + 0.6 × 0.5 = 1.9; a one-number guess of 0.4 against the true 0.5 scores 0.01; the two-number example scores ‖(0.5, −1.0) − (0.4, −0.8)‖² = 0.05. A network that always guesses zero scores 1 per latent number on average, 4,096 for a whole 32 × 32 × 4 latent: the score to beat.

In Python:

import math
z, alpha_bar, eps = 2.0, 0.64, 0.5
# z_t = √ᾱ z + √(1 − ᾱ) ε: the latent, noised to step t in one jump
z_t = math.sqrt(alpha_bar) * z + math.sqrt(1 - alpha_bar) * eps
round(z_t, 2)  # → 1.9
# ‖ε − ε_θ‖² for one number, then for two
round((0.5 - 0.4) ** 2, 2)  # → 0.01
round(sum((e - g) ** 2 for e, g in zip([0.5, -1.0], [0.4, -0.8])), 2)  # → 0.05
# always guessing 0: E[ε²] = 1 per number, over a 32 × 32 × 4 latent
32 * 32 * 4 * 1  # → 4096

What the paper specifies

  • The denoiser εθ(∘, t) is a time-conditional U-Net, “primarily from 2D convolutional layers”, with attention layers at some resolutions.
  • Because ℰ is frozen and the forward process fixed, zt comes straight from ℰ(x) during training, and a finished latent needs one pass through 𝒟 to become a picture.
  • Every model in the appendix tables uses 1,000 diffusion steps and a linear noise schedule; sampling uses DDIM with 50 to 500 steps depending on the experiment.

Why it matters today

Equation 2 is the whole training loop of Stable Diffusion's first versions. The diffusion lesson trains exactly this loss on dots in a plane, and its Step 6 counts the saving of moving to a latent.

In code: add_noise is the one-jump noising and train_noise_predictor trains on this loss; in a latent diffusion model the points they see would be ℰ(x) rather than raw data.

3.3 Conditioning mechanisms · original

“We turn DMs into more flexible conditional image generators by augmenting their underlying UNet backbone with the cross-attention mechanism.”Rombach et al. (2021), §3.3

Everyday picture

A painter working from a brief keeps glancing at it. Each part of the canvas asks its own question of the brief: the sky asks “what colour?”, the foreground asks “what animal?”, and each gets the words that answer it. Cross-attention is that glance: every position in the U-Net's feature map sends a query, and the prompt's tokens answer as keys and values.

Tiny example

Two image positions, three prompt tokens (“a”, “red”, “fox”), width d = 2. The image positions' queries are q1 = (1, 0) and q2 = (0, 1). The encoder τθ gives token vectors a = (0, 0), red = (1, 0), fox = (0, 1); a key projection that doubles them gives keys (0, 0), (2, 0), (0, 2); the value projection leaves them as they are. Position 1 scores the tokens (0, 2, 0)/√2 = (0, 1.414, 0); softmax turns that into shares (0.164, 0.673, 0.164), so position 1 reads mostly “red”: its output is (0.673, 0.164). Position 2 reads mostly “fox”.

φ_i(z_t): image τ_θ(y): prompt W_Q → Q W_K → K W_V → V Q Kᵀ / √d softmax shares · V one output rowper image position

Hover or tap a step. Start with the two inputs at the bottom: the image on the left, the prompt on the right.

One cross-attention layer inside the U-Net, as described in §3.3 and drawn in the paper's Figure 3 as the Q / KV blocks.

Reading it: read from the bottom up. The only difference from ordinary self-attention is where the two inputs come from: the left column starts from the U-Net's own feature map (N image positions), the right from the prompt encoder's output (M tokens). The queries are made from the image; the keys and values from the prompt. The score grid is therefore N × M: every image position against every prompt token. Softmax turns each row into shares, and the shares blend the prompt's values, so each image position receives a mix of prompt information chosen by its own question. The output has one row per image position, the same count as went in, so the layer slots into the U-Net without changing its shape.

In words: “project the image features to queries and the prompt's vectors to keys and values; score every query against every key, shrink the scores by √d, turn each row into shares, and blend the values with them.”

With the numbers: position 1's scores are (0, 2, 0)/√2 = (0, 1.414, 0); softmax gives (0.164, 0.673, 0.164); blending the values (0, 0), (1, 0), (0, 1) gives (0.673, 0.164). Position 2 gives (0.164, 0.673). A note on the paper's shapes: it lists WV as d × dεi and WQ as d × dτ, the reverse of what the formulas need; the query projection must match the image features' width dεi, and the value projection the prompt's width dτ. The Symbols table gives the shapes that fit.

In Python:

import math
# two image positions' queries (W_Q already applied), width d = 2
Q = [[1, 0], [0, 1]]
# τ_θ(y) for "a", "red", "fox"; W_K doubles each vector, W_V keeps it
tau = [[0, 0], [1, 0], [0, 1]]
K = [[2 * v for v in row] for row in tau]
V = tau
K  # → [[0, 0], [2, 0], [0, 2]]
d = 2
for q in Q:
    # Q Kᵀ / √d: this position's score for each token
    scores = [sum(a * b for a, b in zip(q, k)) / math.sqrt(d) for k in K]
    # softmax, then the shares blend the values
    exps = [math.exp(s) for s in scores]
    shares = [e / sum(exps) for e in exps]
    out = [sum(s * v[j] for s, v in zip(shares, V)) for j in range(2)]
    print([round(s, 3) for s in shares], [round(o, 3) for o in out])  # → [0.164, 0.673, 0.164] [0.673, 0.164] [0.164, 0.164, 0.673] [0.164, 0.673]

With the conditioning in place, the prompt encoder τθ and the denoiser εθ are trained together on image-and-condition pairs:

In words: “the same noise-guessing loss, except the network also reads the encoded prompt; its error trains both the denoiser and the prompt encoder.”

With the numbers (illustrative): for the latent above with the prompt “a red fox”, suppose reading the prompt improves the guess from (0.4, −0.8) to (0.45, −0.95). The squared miss drops from 0.05 to 0.005, and the gradient of that miss flows back through the cross-attention into τθ, teaching it which prompt features help.

In Python:

eps = [0.5, -1.0]
# ‖ε − ε_θ(z_t, t)‖² without the prompt, then with it (illustrative guesses)
without = [0.4, -0.8]
with_prompt = [0.45, -0.95]
round(sum((e - g) ** 2 for e, g in zip(eps, without)), 3)  # → 0.05
round(sum((e - g) ** 2 for e, g in zip(eps, with_prompt)), 3)  # → 0.005

Why it matters today

This is the mechanism by which a prompt reaches the pixels in Stable Diffusion and many systems after it. Because τθ can be any encoder of any modality that outputs a sequence of vectors, the same U-Net can be conditioned on text, layouts or class labels without redesign. The paper conditions its class model the same way, with τθ a single embedding table producing one 512-number vector.

In code: scaled_dot_product_attention computes softmax(QKᵀ/√d)V for any queries, keys and values, so passing image features as queries and prompt vectors as keys and values makes it cross-attention; network_inputs is the lesson's simpler conditioning, a label appended to the input.

4 Experiments · original

Everyday picture

The experiments ask two kinds of question. First, the dial: how hard should the autoencoder squeeze? Second, the range: does one recipe serve unconditional pictures, text prompts, class labels, layouts, upscaling and filling holes?

A finding up front

LDMs trained on the VQ-regularized latents sometimes gave better samples than those on KL-regularized latents, even though the VQ autoencoders reconstruct slightly worse (Table 8). Reconstruction quality is not the only property of a latent space that matters to the model trained on it; how easy the latent is to model matters too.

Why it matters

Every model in this section trains on one NVIDIA A100 card (the inpainting model on eight V100s), which is the practical meaning of the paper's claim to democratize high-resolution synthesis.

4.1 On perceptual compression trade-offs · original

Everyday picture

Squeeze too little and the diffusion model still does the autoencoder's job, slowly. Squeeze too much and the autoencoder throws away detail no generator can put back. Somewhere in between is a sweet spot.

Tiny example

The paper names models LDM-f: LDM-1 is plain pixel-space diffusion. For a 256 × 256 image, LDM-1's U-Net works on 65,536 positions per channel map; LDM-4 on 4,096; LDM-8 on 1,024; LDM-32 on 64. Convolution work scales with the number of positions, so LDM-8's feature maps are 64 times smaller than LDM-1's at the same channel count.

In Python:

side = 256
positions = {f: (side // f) ** 2 for f in (1, 2, 4, 8, 16, 32)}
positions  # → {1: 65536, 2: 16384, 4: 4096, 8: 1024, 16: 256, 32: 64}
positions[1] // positions[8]  # → 64

What the paper found

  • With compute fixed (one A100, the same number of steps and parameters), class-conditional ImageNet models with f = 1 or 2 train slowly, and f = 32 stalls early at poor quality. After 2 million steps, LDM-8's FID is 38 points better than LDM-1's (Figure 6 of the paper).
  • Plotting quality against sampling speed for 10 to 200 DDIM steps (Figure 7), LDM-4 and LDM-8 give the best trade-off; pixel-based LDM-1 is both slower and worse.
  • Complex data (ImageNet) prefers milder compression than faces (CelebA-HQ).

Why it matters today

“f = 4 to 16, and 8 is a safe default” is the paper's practical legacy. The same trade-off reappears whenever a new latent space is designed for images or video: every extra factor of compression saves compute and costs fidelity.

4.2 Image generation with latent diffusion · original

Everyday picture

First the plain test: no prompt, just “make a face” or “make a bedroom”, on standard 256 × 256 datasets. Quality is scored by FID, and coverage by a precision and recall for generators: precision asks how many samples look real, recall how much of the real variety the samples cover.

Tiny example

LSUN-Bedrooms: the pixel-space ADM model trained for 232 V100 days with 552 million parameters reaches FID 1.90. LDM-4 uses 274 million parameters and 60 V100 days (by the paper's conversion of one A100 day to 2.2 V100 days) for FID 2.95. That is about half the parameters and about a quarter of the training compute.

In Python:

# ADM against LDM-4 on LSUN-Bedrooms (Table 18)
round(232 / 60, 1)  # → 3.9
round(552 / 274, 1)  # → 2.0
Selected rows of Table 1 of Rombach et al. (2021), unconditional 256 × 256 generation, reproduced with attribution. “N-s” is N DDIM sampling steps; ∗ trained on a KL-regularized latent
DatasetModelFID ↓Prec. ↑Recall ↑
CelebA-HQLSGM7.22
CelebA-HQLDM-4 (500-s)5.110.720.49
FFHQStyleGAN4.160.710.46
FFHQLDM-4 (200-s)4.980.730.50
LSUN-ChurchesDDPM7.89
LSUN-ChurchesLDM-8∗ (200-s)4.020.640.52
LSUN-BedroomsADM1.900.660.51
LSUN-BedroomsLDM-4 (200-s)2.950.660.48

What the rows say

On CelebA-HQ, 5.11 was a new state of the art, ahead of LSGM (which trains a latent diffusion model jointly with its autoencoder). LDMs beat earlier diffusion models everywhere except LSUN-Bedrooms, where they come close to ADM at a fraction of the cost, and the paper notes they consistently improve on GAN-based methods in precision and recall, the coverage a likelihood-based objective is expected to give.

Why it matters today

Matching pixel-space quality at a quarter of the compute is what made the next steps affordable: a 1.45-billion-parameter text model trained by an academic group.

4.3 Conditional latent diffusion · original

Everyday picture

Same painter, new briefs: a sentence, a page of boxes labelled “dog” and “sofa”, a class name. Only the brief reader τθ changes.

4.3.1 Transformer encoders for text and layouts

For text, τθ is a transformer trained jointly with the denoiser: a BERT tokenizer, 77 tokens, 32 layers of width 1,280 (Table 17). The model, LDM-KL-8, has 1.45 billion parameters and trains on LAION-400M image-and-caption pairs. For layouts, each bounding box is encoded as a (top-left, bottom-right, class) token. Samples use classifier-free guidance, which asks the network twice, with and without the prompt, and pushes past the prompted answer; the DiT companion decodes its formula.

Tiny example

Text-to-image on MS-COCO: without guidance the model's FID is 23.31; with guidance scale s = 1.5, it drops to 12.63, alongside GLIDE (12.24, 6 billion parameters) and Make-A-Scene (11.84, 4 billion) with a quarter to a third of their parameters.

In Python:

# parameters, in billions: LDM-KL-8-G against GLIDE and Make-A-Scene
ldm, glide, mas = 1.45, 6, 4
round(ldm / glide, 2), round(ldm / mas, 2)  # → (0.24, 0.36)
# what guidance did to FID
round(23.31 / 12.63, 2)  # → 1.85
Selected rows of Tables 2 and 3 of Rombach et al. (2021), reproduced with attribution
TaskModelFID ↓IS ↑Params
Text-to-image, MS-COCO 256²GLIDE12.246B
Text-to-image, MS-COCO 256²LDM-KL-8, no guidance23.3120.031.45B
Text-to-image, MS-COCO 256²LDM-KL-8-G, s = 1.512.6330.291.45B
Class-conditional ImageNet 256²ADM-G4.59186.7608M
Class-conditional ImageNet 256²LDM-4-G, s = 1.53.60247.67400M

On class-conditional ImageNet, guided LDM-4 reached FID 3.60, better than the pixel-space ADM-G's 4.59 with two thirds of its parameters. That 3.60 is the number the DiT paper later set out to beat. The Inception score (IS) column rises with guidance too.

Why it matters today

This section is the first public demonstration that a mid-sized latent diffusion model with a learned text encoder can compete with far larger text-to-image systems. Later systems commonly use a frozen, pretrained text encoder in τθ's place (the diffusion lesson's pipeline shows CLIP's text tower in that role); the cross-attention wiring is the same.

In code: guided_noise is classifier-free guidance, and guidance_sweep measures what raising the scale does to accuracy and spread.

4.3.2 Convolutional sampling beyond 256², and the scale of the latent · original

Everyday picture

A convolution is a small stencil slid over a grid; it does not care how big the grid is. So a U-Net trained on 256 × 256 pictures can be run on a wider canvas, and when the prompt is itself a picture lined up with the output (a semantic map of “sky here, lake there”), the model keeps things coherent across the larger canvas. The paper generates landscapes up to 1024 × 384 this way.

How the latent is scaled turns out to matter. Noise is added at a fixed size. If the latent's own numbers are large, the same noise hides them less: the signal-to-noise ratio stays high for longer, and the model settles the picture's content early in the reverse process. The KL-regularized latents have a large spread, so the paper rescales them to unit spread before diffusion (Appendices D.1 and G); the VQ latents already have a spread close to 1.

Tiny example

Four latent numbers from a first batch, z = (3, −1, 5, 1). Their mean is 2; their squared distances from it are 1, 9, 9 and 1, averaging 5, so the spread is σ̂ = √5 = 2.236. Dividing by it gives (1.342, −0.447, 2.236, 0.447), which has spread 1.

In words: “over every number in the first batch of latents, find the average and the average squared distance from it; then divide every latent by the square root of that, so the latents have spread 1.”

With the numbers: μ̂ = (3 − 1 + 5 + 1)/4 = 2, σ̂² = (1 + 9 + 9 + 1)/4 = 5, σ̂ = 2.236, and the rescaled latents are (1.342, −0.447, 2.236, 0.447).

In Python:

import math
# a first "batch": four latent numbers
z = [3, -1, 5, 1]
mu_hat = sum(z) / len(z)
sigma2_hat = sum((v - mu_hat) ** 2 for v in z) / len(z)
mu_hat, sigma2_hat  # → (2.0, 5.0)
sigma_hat = math.sqrt(sigma2_hat)
# z ← z / σ̂
[round(v / sigma_hat, 3) for v in z]  # → [1.342, -0.447, 2.236, 0.447]

Reading it: the x-axis is the diffusion step, from 1 (almost clean) to 1,000 (pure noise); the y-axis is the signal-to-noise ratio, Var(z)·ᾱt/(1 − ᾱt), on a base-10 logarithmic scale, so each gridline is a factor of ten. The dashed line is a ratio of 1: signal and noise equally strong. With a unit-spread latent (the slider at 1), the curve crosses 1 near step 260. Drag the spread up to 5, as for an unscaled latent, and the ratio at every step rises 25-fold: the crossing moves to step 567, so over most of the reverse process the content is already legible to the model. That is the paper's “allocates a lot of semantic detail early on” in one curve. The schedule is the linear one from 10−4 to 0.02 over 1,000 steps that the DDPM and DiT papers state; this paper's tables say only “linear”. Computed live from the formula.

Why it matters today

Rescaling the latent by a measured constant is a standard step in latent diffusion pipelines. Whenever a new autoencoder is paired with a diffusion model, this spread must be measured and fixed, or the noise schedule silently means something different.

4.4 Super-resolution with latent diffusion · original

Everyday picture

Handed a thumbnail, a restorer can paint a plausible full-size picture: not the true original, which is lost, but one consistent with the thumbnail. Super-resolution with a diffusion model does exactly that, by feeding the low-resolution image in as a condition (concatenated onto the latent, so τθ is just the identity).

Tiny example

The benchmark upscales ImageNet from 64 × 64 to 256 × 256: 4 times per side, 16 times the pixels. Fifteen of every sixteen output pixels are invented.

In Python:

low, high = 64, 256
# pixels gained, and the share of the output that is new
(high * high) // (low * low)  # → 16
round(1 - (low * low) / (high * high), 4)  # → 0.9375

What the paper found

  • LDM-SR (f = 4, 169 million parameters, 100 steps) scores FID 2.8 against SR3's 5.2 (625 million parameters); SR3 has the better Inception score (180.1 against 166.3) and PSNR (26.4 against 24.4).
  • A plain regression model gets the best PSNR and SSIM of all, which the paper takes as evidence those metrics favour blur over sharp detail that is slightly misplaced.
  • In a user study, people preferred LDM-4's upscales to a pixel-space diffusion baseline's 70.6% of the time.
  • A second model, LDM-BSR, is trained on a mix of realistic degradations (compression, camera noise, blur) and generalizes to real photos, which the bicubic-only model does not.

Why it matters today

Diffusion upscalers conditioned by concatenation are still how many generators reach high resolutions: generate small, then upscale with a second latent diffusion model.

4.5 Inpainting with latent diffusion · original

Everyday picture

Inpainting is filling a hole: a tourist removed from a landmark photo, a scratch mended. A generative model can give several different, equally plausible fillings, where a regression model gives one averaged, often blurry answer.

Tiny example: the speed-up

The paper compares pixel-space LDM-1 with LDM-4 at a fixed parameter count. Training throughput goes from 0.11 to 0.32 samples per second, sampling at 512² from 0.07 to 0.34, and hours per epoch from 20.66 to 7.66; FID after six epochs improves from 24.74 to 15.21. The smallest of those ratios is 2.7, hence the paper's “at least 2.7×” speed-up and “at least 1.6×” better FID.

In Python:

# LDM-1 against LDM-4 (KL, with attention), Table 6
speedups = [0.32 / 0.11, 0.97 / 0.26, 0.34 / 0.07, 20.66 / 7.66]
[round(s, 2) for s in speedups]  # → [2.91, 3.73, 4.86, 2.7]
round(min(speedups), 1)  # → 2.7
round(24.74 / 15.21, 1)  # → 1.6

Reading it: each pair of bars is one measurement from the paper's Table 6, drawn to its own scale so the two models can be compared: the plain bar is pixel-space LDM-1, the striped bar is LDM-4 with a KL-regularized first stage. For the three throughputs longer is better; for hours per epoch and FID shorter is better. On every row the latent model wins by a factor between 1.6 and 4.9. Selected rows, reproduced with attribution.

Selected rows of Table 7 of Rombach et al. (2021): inpainting on 512 × 512 crops of Places, reproduced with attribution
ModelFID, 40 to 50% masked ↓FID, all ↓
LaMa (recomputed on the paper's test set)12.312.23
CoModGAN10.41.82
LDM-4, big, fine-tuned at 512²9.391.50

Why it matters today

The big fine-tuned model set a new state of the art on the benchmark, and people preferred its fillings to LaMa's in a user study (68.1% in a head-to-head choice). The fine-tuning detail is worth keeping: its samples at 512² were worse than at 256², which the authors attribute to its added attention layers, and half an epoch of training at 512² let it adjust.

5 Limitations and societal impact · original

Limitations

Sampling is still a loop of many network calls, slower than a GAN's single pass. And the autoencoder is a bottleneck for pixel-exact work: whatever it cannot reconstruct, the system cannot produce. Even the mild f = 4 autoencoders lose a little, which the authors suspect already limits their super-resolution.

Tiny example

With PSNR's usual definition on a 0 to 255 scale, the f = 4 VQ autoencoder's 27.43 dB is a root-mean-squared error of about 11 grey levels: invisible in a photo, visible in a task that needs exact pixels, such as recovering small text.

In Python:

# PSNR = 20 log10(255 / RMSE), so RMSE = 255 / 10^(PSNR / 20)
psnr = 27.43
round(255 / 10 ** (psnr / 20), 1)  # → 10.8

Societal impact

The paper names both sides: cheaper generation broadens creative use and access, and also makes manipulated images, misinformation and non-consensual “deep fakes” easier to produce, with women disproportionately harmed. Generative models can reveal training data, and they can reproduce or amplify dataset biases; how much the two-stage design misrepresents the data is left as an open question.

Why it matters today

Releasing the models (contribution vi) is what made both sides of that paragraph real within a year.

6 Conclusion · original

What the paper concludes

Latent diffusion models improve both training and sampling efficiency of denoising diffusion models without degrading quality, and with cross-attention conditioning they compete with the best task-specific methods across a range of conditional image tasks, with no task-specific architecture.

Why it matters today

The two ideas separate cleanly, which is why both survived: the latent space is now standard for image and video diffusion, and cross-attention (or its successors) is how prompts reach the denoiser.

Appendix B: the objective, through signal-to-noise · original

Everyday picture

Different papers describe the same noise schedule with different letters. The appendix uses one number per step that everyone can agree on: the signal-to-noise ratio, how loud the picture is compared with the static on top of it. The ratio starts huge and falls towards zero, and each step of the training bound turns out to be weighted by how much the ratio falls at that step.

Tiny example

Use the two-step toy chain from the DDPM companion: the signal left after one step is ᾱ1 = 0.9 and after two ᾱ2 = 0.72. A warning about letters: this appendix writes αt for the signal scale, which is √ᾱt in DDPM's notation, and σt for the noise scale √(1 − ᾱt). So SNR(1) = 0.9/0.1 = 9 and SNR(2) = 0.72/0.28 = 2.571.

In words: “at step t the noisy picture is the clean one scaled by αt plus noise of spread σt; the signal-to-noise ratio is the square of the one over the square of the other.”

With the numbers: α1 = √0.9, σ1 = √0.1, so SNR(1) = 9; α2 = √0.72, σ2 = √0.28, so SNR(2) = 2.571.

In Python:

import math
alpha_bar = {1: 0.9, 2: 0.72}
def snr(t):
    # α_t = √ᾱ_t and σ_t = √(1 − ᾱ_t) in this appendix's letters
    alpha_t, sigma_t = math.sqrt(alpha_bar[t]), math.sqrt(1 - alpha_bar[t])
    return alpha_t ** 2 / sigma_t ** 2
round(snr(1), 3), round(snr(2), 3)  # → (9.0, 2.571)

If the model's reverse step is built from the true backward step with the clean picture x0 replaced by a guess xθ, the bound's middle terms become weighted squared errors, and the noise-guessing form follows by one substitution:

In words: “each step's share of the bound is half the drop in signal-to-noise ratio at that step, times the squared error of the guessed clean picture; and an error in the guessed clean picture is 1/SNR times the same error measured in noise units. Dropping all those weights, and counting every step equally, gives equation 1.”

With the numbers: at t = 2, the weight on the clean-picture error is ½(9 − 2.571) = 3.214, and the conversion to noise units is σ²/α² = 0.28/0.72 = 0.389, so the weight on the noise error is 3.214 × 0.389 = 1.25. As a check, this matches the DDPM paper's own weight βt²/(2σt²αt(1 − ᾱt)) with its variance choice β̃t: 0.2 / (2 × 0.8 × 0.1) = 1.25. Two small slips in the appendix's printed equation 13: the KL's second argument is written p(xt−1) (it should be p(xt−1|xt)), and a closing bracket is missing; both are fixed above.

In Python:

alpha_bar = {1: 0.9, 2: 0.72}
snr = {t: a / (1 - a) for t, a in alpha_bar.items()}
# ½(SNR(t − 1) − SNR(t)) at t = 2
w_x0 = 0.5 * (snr[1] - snr[2])
round(w_x0, 3)  # → 3.214
# σ_t² / α_t²: from clean-picture error to noise error
to_eps = (1 - alpha_bar[2]) / alpha_bar[2]
round(to_eps, 3)  # → 0.389
round(w_x0 * to_eps, 4)  # → 1.25
# DDPM's weight β / (2 α (1 − ᾱ_(t−1))) with σ_t² = β̃_t: the same number
beta, alpha = 0.2, 0.8
round(beta / (2 * alpha * (1 - alpha_bar[1])), 4)  # → 1.25

Why it matters today

Writing schedules through the signal-to-noise ratio is how later work compares and designs them: two schedules with the same SNR curve define the same diffusion. It also explains §4.3.2 in one line: scaling the latent by s multiplies every SNR by s².

Appendix C: guiding with an image · original

Everyday picture

An unconditional model can still be steered while it samples. At each step, look at the picture the model is heading for, measure how far it is from what you want (for example, “when shrunk, it should look like this thumbnail”), and nudge the noise guess to reduce that distance. The paper uses this to upscale unconditional samples and to push up super-resolution PSNR.

Tiny example

One latent number: the network's noise guess is εθ = 0.5 at a step where σt = √(1 − αt²) = 0.6. The decoded, shrunk guess of the final picture has one pixel at 0.6, and the target thumbnail says 1.0. The guider's log-probability is −½(1.0 − 0.6)² = −0.08, and its slope with respect to that pixel is 1.0 − 0.6 = 0.4. Suppose (illustratively) backpropagation carries a slope of 0.4 back to zt.

In words: “predict the finished latent, decode it, transform it the way the target was made, and measure the squared distance to the target; then correct the noise guess by the slope of that measurement, scaled by the current noise level.”

With the numbers: log pΦ = −½ × 0.4² = −0.08. As printed, ε̂ = 0.5 + 0.6 × 0.4 = 0.74. A caution on the sign: since the score is the noise guess flipped and divided by σt, adding ∇log pΦ to the score means subtracting σt∇log pΦ from the noise guess, giving 0.5 − 0.24 = 0.26. The paper prints a plus sign; the derivation gives a minus, so read equation 16 as ε̂ ← εθ − σt∇log pΦ.

In Python:

eps_theta, sigma_t = 0.5, 0.6
y, decoded = 1.0, 0.6
# log p_Φ(y | z_t) = −½ ‖y − T(D(z_0(z_t)))‖², and its slope for the decoded pixel
log_p = -0.5 * (y - decoded) ** 2
round(log_p, 2)  # → -0.08
slope = y - decoded
# the correction as printed, and with the sign the score relation gives
round(eps_theta + sigma_t * slope, 2), round(eps_theta - sigma_t * slope, 2)  # → (0.74, 0.26)

Why it matters today

“Guide sampling with the gradient of any differentiable measurement of the predicted image” became a general tool: guiding by a classifier, by a CLIP score, by a thumbnail, by a mask. Classifier-free guidance (§4.3.1) replaced it for prompts, because it needs no extra network and no backpropagation at every step.

What changed since 2021

The two-stage design is the part that stayed. Around it:

Choice in the paperCommon todayWhyWhere to read
Convolutional U-Net denoiser with cross-attentionOften a transformer over latent patchesScales predictably with computeDiT companion
Text encoder trained jointly with the denoiserFrozen pretrained text encodersBetter language understanding for freeCLIP companion
Noise prediction, 50 to 250 DDIM stepsAlso flow matching, and distilled few-step samplersFewer network calls per imageFlow matching companion
ImagesVideo in a latent space with a time axisThe same compression argument, with far more pixelsmultimodal lesson

Glossary

Every term with hover guidance on this page, in one place.