Flow Matching for Generative Modeling, annotated
How to read this page
Nothing here assumes you already know the jargon. Three things help:
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation is followed by a table of its symbols, a sentence reading it aloud, the numbers of a tiny example, and the same example in a few lines of Python.
- The pictures are live: drag the sliders, click (or use the arrow keys) to move points, and watch the readouts.
One warning about labels, which this paper shares with the diffusion and flow matching lesson: here x0 is the noise, x1 is the data, and time runs from 0 (noise) to 1 (data). Diffusion papers use the opposite convention.
Abstract
“Specifically, we present the notion of Flow Matching (FM), a simulation-free approach for training CNFs based on regressing vector fields of fixed conditional probability paths.”Lipman et al. (2022), Abstract. Read the original
Everyday picture
Picture a weather map covered in little arrows, one at every place, saying which way the wind blows and how hard. Drop a leaf anywhere and it drifts along the arrows. Now suppose you want a wind that carries leaves scattered at random (noise) into the shape of, say, a face (data). This paper is a recipe for learning that wind. The wind is a velocity field, and a generator built by following one is a continuous normalizing flow.
What the paper claims
- A way to train these flows with no simulation during training: each step is one network call and a squared error, like training a diffusion model.
- The recipe works with a whole family of paths from noise to data. Diffusion's paths are members of the family, and training them this way was more stable than the usual score matching.
- A new member, optimal transport paths, moves every point in a straight line at constant speed. It trained faster, sampled in fewer steps and scored better on ImageNet than the diffusion-based methods it was compared with.
Why it matters today
Flow matching with straight paths, and its close cousin rectified flow, is how several of today's image generators are trained, Stable Diffusion 3 among them. The diffusion lesson builds it from scratch in a few dozen lines of NumPy.
1 Introduction · original
“We find that conditional OT paths are simpler than diffusion paths, forming straight line trajectories whereas diffusion paths result in curved paths.”Lipman et al. (2022), §1
Everyday picture
Continuous flows were an elegant idea with a painful training bill. To check whether the wind was right, you had to actually release leaves and follow them all the way, step by step, then trace back through every step to work out how to adjust the wind. Diffusion models escaped that bill with a trick: they never follow the leaves during training, they only ask the network one question about one moment. The paper's goal is to give general flows the same trick.
Tiny example
Suppose (illustratively) that following the wind accurately from noise to data takes 100 small steps. Training by maximum likelihood then costs about 100 network calls forward, plus the backward pass through all of them, for every training example. Flow matching costs one call per example: pick a moment, look at one point, grade one guess.
Why it matters
“Simulation-free” is the whole game. It is what let diffusion models scale to billions of images, and this paper extends it to any path you can write down, including straighter ones that are cheaper to follow when you generate.
2 Preliminaries: continuous normalizing flows · original
“A CNF is used to reshape a simple prior density p0 (e.g., pure noise) to a more complicated one, p1, via the push-forward equation”Lipman et al. (2022), §2
Everyday picture
Two objects carry the whole paper. A velocity field is the wind map: at each place and moment, an arrow. A probability path is a movie of a crowd: frame 0 is a cloud of random noise, frame 1 is the data, and every frame in between is a probability distribution saying where the crowd is at that moment. Let every person in the frame-0 crowd walk with the wind, and you get a movie. The wind generates that movie.
Tiny example
Take a one-dimensional wind that blows outwards in proportion to where you stand: v(x) = x. A walker starting at x = 1 moves at speed 1, then faster as they get further out, and after time ln 2 ≈ 0.693 has reached exactly 2. Every walker doubles their position in that time, so the whole crowd spreads out to twice its width, and the height of its bell curve halves (the same number of people over twice the ground).
The math
The flow φ records where a walker who started at x is at time t. It is defined by an ordinary differential equation: the walker's velocity at every moment is whatever the wind says at their current spot.
In words: “a walker's position changes at the rate the velocity field gives at that position and time, and at time 0 every walker is where they started.”
With the numbers: with v(x) = x, the walker from x = 1 is at φt(1) = et, whose rate of change is et, its own position, as the equation demands. At t = ln 2 it is at 2. Following the field in 100,000 tiny Euler steps lands on 2.0 too.
In Python:
import math
# the exact flow for v(x) = x: φ_t(x) = x e^t
t = math.log(2)
1 * math.exp(t) # → 2.0
# the same, by following the field in small Euler steps
x, n = 1.0, 100_000
h = t / n
for _ in range(n):
x = x + h * x
round(x, 3) # → 2.0
The second equation says how the crowd's density changes as it moves. Where walkers spread apart, the density drops, by exactly how much the flow stretches space there.
In words: “to find the crowd's density at x at time t, run the flow backwards to see where the walkers now at x started, read the starting density there, and correct for how much the flow stretched or squeezed the space around them.”
With the numbers: with v(x) = x, φt−1(x) = x e−t and the stretch factor is e−t. Start from the standard bell curve, whose height at 0 is 0.3989. At t = ln 2 the height at 0 is 0.3989 × 0.5 = 0.1995, and at x = 2 it is p0(1) × 0.5 = 0.2420 × 0.5 = 0.121.
In Python:
import math
# p_0: the standard bell curve
def p0(x):
return math.exp(-x * x / 2) / math.sqrt(2 * math.pi)
t = math.log(2)
# φ_t^-1(x) = x e^-t, and det[∂φ_t^-1/∂x] = e^-t
def p_t(x):
return p0(x * math.exp(-t)) * math.exp(-t)
round(p_t(0), 4) # → 0.1995
round(p_t(2), 4) # → 0.121
Hover or tap the chart, or focus it and use the arrow keys, to read the two curves.
Reading it: the x-axis is position and the y-axis is density (how crowded that spot is). The solid curve is the crowd at time 0, the standard bell curve. The dashed curve is the same crowd after walking with the wind v(x) = x for time t, computed from the push-forward formula above. Drag t up and the crowd spreads out: the curve gets wider and lower, because the same people now cover more ground. The area under both curves stays exactly 1; nobody is created or lost. At t = 0.69 (ln 2) the dashed peak is half the solid one.
Why it matters today
A continuous normalizing flow is exactly this machine, with a neural network vt(x; θ) as the wind. Generating means solving the first equation from a noise sample; measuring how probable a data point is means running the second one. Diffusion models, sampled deterministically, are also flows of this kind, which is why the diffusion lesson can treat DDIM and flow matching side by side. Its euler_step is the loop in the Python above.
3 Flow Matching · original
“Simply put, the FM loss regresses the vector field ut with a neural network vt.”Lipman et al. (2022), §3
Everyday picture
Suppose someone hands you a perfect wind map, u, drawn by hand. You could train a network to copy it the way a student copies a map: at random places and times, compare the network's arrow with the true arrow, and nudge the network to shrink the difference. That is the Flow Matching objective. It is plain regression: predict a number (here, an arrow) and be graded by squared error.
Tiny example
At some place and time the true arrow says “move right at speed 3”, and the network says 2.5. The loss there is (2.5 − 3)2 = 0.25. Average this over many random places and times and you have the loss.
In words: “pick a random time and a random point from the crowd at that time; measure how far the network's arrow is from the true arrow, squared; average over many picks.”
With the numbers: one pick with true arrow 3 and guess 2.5 contributes (2.5 − 3)2 = 0.25.
In Python:
# one pick: the true arrow u_t(x) and the network's guess v_t(x)
u, v = 3.0, 2.5
# ‖v − u‖²
(v - u) ** 2 # → 0.25
The catch
Nobody has the perfect map. To use this loss you would need to know, for real data such as photographs, the exact probability path from noise to photos and the exact wind that produces it. Both are out of reach. The rest of Section 3 is the paper's way around this, and it is the heart of the paper.
3.1 Building the path, and its wind, from one-example pieces · original
“The marginal vector field (equation 8) generates the marginal probability path (equation 6).”Lipman et al. (2022), §3.1
Everyday picture
Think of a railway network. For a single destination, the timetable is easy to write: one train, one track, from the depot (noise) to that one town (a data point). The paper builds the whole network out of these one-destination timetables. The crowd at any moment is simply everybody on every train, a mixture. And the wind at a spot is the average of the directions of the trains passing through it, weighted by how many of each are there.
Tiny example
Take a one-dimensional dataset with just two points, x1 = −2 and x1 = +2, equally likely. For each destination, use the straight track of the lesson: at time t its small crowd is a bell curve centred on t·x1 with spread 1 − t, and a rider at x travels at (x1 − x)/(1 − t): the distance still to go, divided by the time left (§4 derives this). Stand at x = 0.5 at time t = 0.5.
- The “+2” crowd is centred at 1 with spread 0.5: its density at 0.5 is 0.4839. The “−2” crowd is centred at −1: its density at 0.5 is only 0.0089.
- So of the riders passing 0.5, a share 0.982 are heading to +2 and 0.018 to −2.
- Riders to +2 need (2 − 0.5)/0.5 = 3; riders to −2 need (−2 − 0.5)/0.5 = −5. The wind at this spot is the weighted average: 0.982 × 3 + 0.018 × (−5) ≈ 2.86.
The math
The whole crowd is the mixture of the one-destination crowds, weighted by how common each data point is:
In words: “the density of the whole crowd at x is the density of each data point's small crowd at x, weighted by how common that data point is, summed over all data points.”
With the numbers: p0.5(0.5) = ½ × 0.0089 + ½ × 0.4839 = 0.2464. At t = 1 each small crowd shrinks onto its data point, so the mixture becomes (almost exactly) the data itself: that is the paper's equation 7.
In Python:
import math
# a bell curve with mean mu and spread s, read at x
def N(x, mu, s):
return math.exp(-(x - mu) ** 2 / (2 * s * s)) / (s * math.sqrt(2 * math.pi))
t, x = 0.5, 0.5
data = [-2.0, 2.0]
q = [0.5, 0.5]
# p_t(x | x_1): each data point's crowd, centred at t·x_1 with spread 1 − t
p_cond = [N(x, t * x1, 1 - t) for x1 in data]
[round(p, 4) for p in p_cond] # → [0.0089, 0.4839]
# ∫ ... q(x_1) dx_1 becomes a sum over the two data points
p_t = sum(p * q_k for p, q_k in zip(p_cond, q))
round(p_t, 4) # → 0.2464
The wind of the whole crowd is the average of the one-destination winds, weighted by each destination's share of the riders at x:
In words: “the wind at x is each data point's wind at x, weighted by the share of the crowd at x that belongs to that data point.”
With the numbers: the shares are 0.0089 × 0.5 / 0.2464 = 0.018 and 0.4839 × 0.5 / 0.2464 = 0.982; the winds are −5 and 3; so u0.5(0.5) = 0.018 × (−5) + 0.982 × 3 = 2.86.
In Python:
import math
def N(x, mu, s):
return math.exp(-(x - mu) ** 2 / (2 * s * s)) / (s * math.sqrt(2 * math.pi))
t, x = 0.5, 0.5
data, q = [-2.0, 2.0], [0.5, 0.5]
p_cond = [N(x, t * x1, 1 - t) for x1 in data]
p_t = sum(p * q_k for p, q_k in zip(p_cond, q))
# p_t(x | x_1) q(x_1) / p_t(x): each data point's share of the crowd at x
shares = [p * q_k / p_t for p, q_k in zip(p_cond, q)]
[round(s, 3) for s in shares] # → [0.018, 0.982]
# u_t(x | x_1): distance still to go over time left
u_cond = [(x1 - x) / (1 - t) for x1 in data]
u_cond # → [-5.0, 3.0]
# the weighted average of the one-destination winds
u_t = sum(s * u for s, u in zip(shares, u_cond))
round(u_t, 2) # → 2.86
See it in two dimensions
Click the picture, or focus it and use the arrow keys, to move the probe point.
Reading it: the four solid squares are the data. The dashed circles are the four one-destination crowds at the chosen time: each is centred on t·x1, and its radius is its spread. The small grey arrows everywhere are the marginal wind ut(x) of equation 8. At the probe point (the ring), the thin arrows are the four one-destination winds, each drawn fainter the smaller its share (the percentages), and the thick red arrow is their weighted average, the marginal wind. Start at t = 0: all four crowds sit on top of each other, every share is 25%, and the thick arrow points at the average of the data, the empty middle. Drag t up and the crowds separate; soon one share takes almost everything, and the thick arrow swings to follow the nearest corner. That is how a single wind, with no randomness at all, splits one cloud of noise into four clusters.
Theorem 1, in words
The surprising part is that this averaged wind really does move the averaged crowd: if each one-destination wind produces its own crowd, the weighted-average wind of equation 8 produces the mixture crowd of equation 6. The proof (Appendix A) checks the continuity equation, which says that mass is neither created nor destroyed. Nothing about images is assumed; it holds for any data distribution.
Why it matters today
This is the move that turns an impossible target into a buildable one: you never need the wind for all of the data at once, only the wind for one data point at a time, which you can write down by hand. The toy picture above even shows why one Euler step from noise lands in the middle of the data, a failure the diffusion lesson measures on its own four blobs.
3.2 Conditional Flow Matching · original
“The FM (equation 5) and CFM (equation 9) objectives have identical gradients w.r.t. θ.”Lipman et al. (2022), §3.2
Everyday picture
Equation 8 still hides an integral over all the data, so even the averaged wind cannot be computed for a real dataset. The paper's answer is to skip the average entirely. Pick one data point, put a rider on its track, and train the network to predict that rider's velocity. Different riders passing the same spot want different arrows, so the network can never be perfect. The best it can do is answer with their average, which is exactly the averaged wind it could not compute. The average takes care of itself.
Tiny example
At x = 0.5, t = 0.5 above, 98.2% of riders want 3 and 1.8% want −5. If the network answers v, its expected squared error is 0.982(v − 3)2 + 0.018(v + 5)2. For v = 2.5 that is 1.26; the error against the true averaged wind would be (2.5 − 2.86)2 = 0.13. The gap, 1.13, is the same whatever v is (try v = 0: 9.29 against 8.16). A gap that never changes cannot change the slope, so both losses pull the network the same way.
In words: “pick a time, pick one real data point, put a point on that data point's path at that time, and grade the network's arrow against that one path's arrow.”
With the numbers: one draw where the rider is heading to +2 grades the guess 2.5 as (2.5 − 3)2 = 0.25; a draw heading to −2 grades it as (2.5 + 5)2 = 56.25. Over many draws at this spot, 98.2% are the first kind, and the average is 1.26.
In Python:
# at x = 0.5, t = 0.5: each destination's share and its wind
shares, u_cond = [0.018, 0.982], [-5.0, 3.0]
v = 2.5
# the CFM loss at this spot: the average of the per-rider errors
cfm = sum(s * (v - u) ** 2 for s, u in zip(shares, u_cond))
round(cfm, 2) # → 1.26
The one line of algebra behind Theorem 2: for a fixed spot, the average squared distance to a set of targets is the squared distance to their average, plus how spread out the targets are.
In words: “the conditional loss at a spot equals the marginal loss at that spot plus a leftover that depends only on the data, not on the network.” Averaging over spots and times turns the left side into the CFM loss and the first term on the right into the FM loss; the leftover is the constant of Theorem 2.
With the numbers: for v = 2.5: 1.26 = 0.13 + 1.13. For v = 0: 9.29 = 8.16 + 1.13. The leftover, 1.13, is the spread of the targets 3 and −5 around their average 2.86.
In Python:
shares, u_cond = [0.01798621, 0.98201379], [-5.0, 3.0]
# u_t(x): the average target
u_t = sum(s * u for s, u in zip(shares, u_cond))
# the leftover: how spread out the targets are around u_t(x)
spread = sum(s * (u - u_t) ** 2 for s, u in zip(shares, u_cond))
round(spread, 2) # → 1.13
for v in (2.5, 0.0):
cfm = sum(s * (v - u) ** 2 for s, u in zip(shares, u_cond))
fm = (v - u_t) ** 2
print(round(cfm, 2), round(fm, 2), round(cfm - fm, 2)) # → 1.26 0.13 1.13 9.29 8.16 1.13
Reading it: the x-axis is the network's answer v at one spot, over a window 3 either side of the best answer; the y-axis is the loss. The solid curve is the Flow Matching loss, graded against the true averaged wind; the dashed curve is the Conditional Flow Matching loss, graded against each rider's own wind. They are the same bowl, one lifted above the other by a constant, so their lowest points sit at the same v and their slopes match everywhere. Slope is what training follows, so training on the computable dashed loss is training on the uncomputable solid one. Drag x to 0: the two destinations are now equally likely, the targets are +4 and −4, the best answer is 0, and the gap grows to 16, because the riders disagree completely.
Why it matters today
This is why the lesson's training loop is so short: train_velocity_predictor pairs each data point with fresh noise, picks a time, and regresses one rider's velocity, exactly the CFM loss. The same argument, with noise in place of velocity, is why diffusion's “guess the noise” loss works: the paper credits denoising score matching as its inspiration.
4 Conditional probability paths and vector fields · original
“We decide to use the simplest vector field corresponding to a canonical transformation for Gaussian distributions.”Lipman et al. (2022), §4
Everyday picture
Now choose the one-destination tracks. The paper uses a bell-shaped crowd (a Gaussian) that slides and shrinks over time: at time 0 it is the standard noise cloud centred on 0, at time 1 a tiny cloud centred on the data point. Two dials describe any such track: where the centre is at each moment (the mean) and how wide the cloud is (the spread). The simplest way to move a bell curve is to stretch it and shift it, and the wind that does exactly that is what Theorem 3 writes down.
Tiny example
Data point x1 = 2; centre 2t; spread 1 − t. A noise sample x0 = −1 is carried to −1 × (1 − t) + 2t: at t = 0.25 that is −0.25. The centre moves at speed 2 and the spread shrinks at rate 1, and together they move this point at speed 3, the same answer as the lesson's straight line from −1 to 2.
The math
In words: “the crowd heading to x1 is, at every moment, a round bell curve with centre μt and spread σt,” with μ0 = 0, σ0 = 1 (pure noise) and μ1 = x1, σ1 = σmin (a tight cloud on the data).
With the numbers: at t = 0.25 the crowd heading to 2 has centre 0.5 and spread 0.75, and its density at −0.25 is 0.3226.
In Python:
import math
x1, t = 2.0, 0.25
mu, sigma = t * x1, 1 - t
mu, sigma # → (0.5, 0.75)
# N(x | μ, σ²) in one dimension, read at x = −0.25
x = -0.25
round(math.exp(-(x - mu) ** 2 / (2 * sigma ** 2)) / (sigma * math.sqrt(2 * math.pi)), 4) # → 0.3226
The flow that produces this crowd from noise stretches and shifts each noise sample:
In words: “to place a noise sample on the track at time t, scale it by the current spread and add the current centre.” Training with it is equation 14 of the paper: draw noise x0, place it with ψt, and grade the network against the speed of ψt(x0).
With the numbers: ψ0.25(−1) = 0.75 × (−1) + 0.5 = −0.25.
In Python:
x0, x1, t = -1.0, 2.0, 0.25
# σ_t x + μ_t
psi = (1 - t) * x0 + t * x1
psi # → -0.25
Theorem 3 gives the wind that moves every point of that crowd along its track:
In words: “a point's velocity is the centre's velocity, plus its offset from the centre scaled by the rate at which the cloud is shrinking or growing.” A point off-centre in a shrinking cloud is pulled inwards; the centre itself just moves with μ′.
With the numbers: at x = −0.25, t = 0.25: σ′ = −1, σ = 0.75, μ = 0.5, μ′ = 2, so u = (−1 / 0.75) × (−0.25 − 0.5) + 2 = 1 + 2 = 3.
In Python:
x1, t, x = 2.0, 0.25, -0.25
# μ_t = t x_1 and σ_t = 1 − t, and their rates of change
mu, mu_prime = t * x1, x1
sigma, sigma_prime = 1 - t, -1.0
u = sigma_prime / sigma * (x - mu) + mu_prime
round(u, 6) # → 3.0
Why it matters today
Choosing μt and σt is now a design decision, not something inherited from a noise process. The next two examples are two choices of these dials: one gives back diffusion, the other gives straight lines.
Example I: diffusion paths · original
“Another important observation is that, as these probability paths were previously derived as solutions of diffusion processes, they do not actually reach a true noise distribution in finite time.”Lipman et al. (2022), §4.1
Everyday picture
Diffusion's tracks were never designed as tracks: they fall out of a process that fades a photo into static a little at a time. Read backwards (noise to data), the fading gives a centre and a spread for every moment, so diffusion is just one setting of the two dials. The paper checks the two classic settings, variance exploding and variance preserving, and finds that Theorem 3 gives the same wind as diffusion's deterministic sampler (the probability flow; Appendix D).
Tiny example
With the paper's settings βmin = 0.1 and βmax = 20 (Appendix E), halfway along (t = 0.5) the variance-preserving centre is only 0.28 of the way to the data, and the spread is still 0.96: the point is nearly all noise. At the start (t = 0) the centre is 0.0066·x1, not 0. The noise at t = 0 is almost the standard cloud, but not quite.
In words: “the variance-preserving track shrinks the data point by α and fills the rest with noise, so that the squares of the data share and the noise share always add to 1; α decays exponentially with the accumulated noise rate T, read backwards in time.”
With the numbers: with β(s) = 0.1 + 19.9s, T(s) = 0.1s + 9.95s2. At t = 0.5: T(0.5) = 2.5375, α0.5 = e−1.26875 = 0.2812, spread √(1 − 0.28122) = 0.9597. At t = 0: α1 = e−5.025 = 0.0066.
In Python:
import math
beta_min, beta_max = 0.1, 20.0
# T(t) = ∫_0^t β(s) ds with β(s) = β_min + s (β_max − β_min)
def T(t):
return t * beta_min + 0.5 * t * t * (beta_max - beta_min)
def alpha(t):
return math.exp(-0.5 * T(t))
t = 0.5
round(T(1 - t), 4), round(alpha(1 - t), 4) # → (2.5375, 0.2812)
# the spread √(1 − α²)
round(math.sqrt(1 - alpha(1 - t) ** 2), 4) # → 0.9597
# at t = 0 the centre is not quite 0
round(alpha(1.0), 4) # → 0.0066
Why it matters today
Two things follow. Training diffusion's own tracks with the Flow Matching loss is a legitimate alternative to score matching, and the paper found it steadier (§6). And because diffusion's tracks never quite reach pure noise, diffusion code has to paper over the gap; flow matching can simply set μ0 = 0 and σ0 = 1 exactly. The lesson's add_noise is this same shrink-and-add-noise step, in the diffusion convention.
Example II: optimal transport paths · original
“Intuitively, particles under the OT displacement map always move in straight line trajectories and with constant speed.”Lipman et al. (2022), §4.1
Everyday picture
The most natural dials of all: move the centre at a steady pace from 0 to the data point, and shrink the spread at a steady pace from 1 to almost nothing. Every noise sample then travels in a straight line at constant speed, like a courier who knows the address and never turns. Moving one crowd onto another with the least total effort is what optimal transport means, and for these two bell curves the straight-line plan is the optimal one.
Tiny example
Noise x0 = −1, data x1 = 2, σmin = 0.01. At t = 0.25 the sample sits at 0.7525 × (−1) + 0.25 × 2 = −0.2525, and its velocity is 2 − 0.99 × (−1) = 2.99 at every moment. With σmin = 0 these are the lesson's −0.25 and 3.
In words: “the centre slides from 0 to the data point at constant speed, and the spread shrinks from 1 to σmin at constant speed.”
With the numbers: at t = 0.25 with x1 = 2: centre 0.5, spread 1 − 0.99 × 0.25 = 0.7525.
In Python:
x1, t, sigma_min = 2.0, 0.25, 0.01
mu = t * x1
sigma = 1 - (1 - sigma_min) * t
mu, sigma # → (0.5, 0.7525)
Plugging these dials into Theorem 3 gives the conditional wind (the paper's equation 21), which is defined all the way to t = 1:
In words: “the velocity is the distance still to cover (to the data point, allowing for the tiny final spread) divided by the time left.” For a fixed x the direction never changes with time, only the size: the paper's Figure 2 contrasts this with diffusion's target, which turns as time passes.
With the numbers: at x = −0.2525, t = 0.25: (2 − 0.99 × (−0.2525)) / 0.7525 = 2.99.
In Python:
x1, t, sigma_min = 2.0, 0.25, 0.01
x = -0.2525
u = (x1 - (1 - sigma_min) * x) / (1 - (1 - sigma_min) * t)
round(u, 4) # → 2.99
The flow itself (equation 22) is a straight line from the noise sample to (almost) the data point:
In words: “mix the noise sample and the data point: the data's share grows from 0 to 1 at a steady pace, and the noise's share shrinks from 1 to σmin.”
With the numbers: (1 − 0.99 × 0.25) × (−1) + 0.25 × 2 = −0.2525.
In Python:
x0, x1, t, sigma_min = -1.0, 2.0, 0.25, 0.01
psi = (1 - (1 - sigma_min) * t) * x0 + t * x1
round(psi, 4) # → -0.2525
And the training loss (equation 23) is the lesson's loss, with σmin kept:
In words: “draw a time, a data point and a noise sample; place the noise on the straight line toward the data; grade the network's velocity there against the line's constant velocity.”
With the numbers: the target is 2 − 0.99 × (−1) = 2.99; a guess of 2.5 scores (2.5 − 2.99)2 = 0.2401.
In Python:
x0, x1, sigma_min = -1.0, 2.0, 0.01
# x_1 − (1 − σ_min) x_0: the same at every t
target = x1 - (1 - sigma_min) * x0
round(target, 4) # → 2.99
round((2.5 - target) ** 2, 4) # → 0.2401
Figure 3, redrawn: two tracks from the same noise
Reading it: the hollow circle is the noise sample x0 and the filled square is the data point x1. The solid line is the optimal transport track: straight, and the round dot moves along it at an even pace as you drag t. The dashed curve is the diffusion track: it bows outwards, because its noise share and data share are the sides of a right triangle (their squares add to 1), and its square dot creeps through the first half (at t = 0.5 it is only 28% data), then rushes to the data at the end. A solver taking a few big straight steps follows the straight track exactly; on the curved, late-rushing track big steps cut the corner and land off it. The paper also sees diffusion's sampling paths “overshoot” the final sample and double back.
Hover or tap the chart, or focus it and use the arrow keys, to read both curves.
Reading it: the x-axis is time from noise (0) to data (1); the y-axis is how much of the moving point is still noise, σt. The solid line is optimal transport: noise drains away at a steady rate. The dashed curve is the variance-preserving diffusion track: at t = 0.5 the point is still 96% noise, and most of the cleaning happens in the last third. That is the pattern the paper sees in real images (its Figure 6): the optimal transport model's pictures emerge gradually, while diffusion's stay static-like until the very end.
The whole recipe in one picture
Hover or tap a block. Start with noise x₀ at the top left.
Reading it: read the top box from the top down. Three random draws (a noise sample, a time and a real data point) feed two calculations: the point on the straight line at that time (left) and the line's velocity (right, which ignores t because the speed is constant). The network sees only the point and the time, guesses a velocity, and the squared difference is the loss. Nothing is simulated: one training step is one network call. The bottom box is generation: start from noise at t = 0 and repeat solver steps along the learned wind until t = 1. The paper uses an adaptive solver (dopri5) or fixed-step ones (Euler, midpoint); the Euler step is shown.
Why it matters today
This is the version everyone uses: the lesson's flow_point and flow_target are equations 22 and 23 with σmin = 0. The paper also makes a careful point that later work took up: each conditional track is straight and optimal, but the averaged wind is not guaranteed to be; its paths bend where tracks cross. The lesson's figure of crossing training lines and bending learned paths shows exactly this, and rectified flow's retraining exists to straighten them.
5 Related work · original
“In contrast, the Flow Matching framework allows simulation-free training with unbiased gradients and readily scales to very high dimensions.”Lipman et al. (2022), §5
Everyday picture
Before this paper there were three kinds of road. Train flows by following them (accurate, but slow and hard to scale). Train them without following, but with sums that are hard to estimate or gradients that are slightly wrong. Or build a diffusion process and let it dictate the track. Flow Matching takes the last road's trick (grade one example at a time, as denoising score matching does) and applies it to any track you like.
What sits nearby
- Maximum-likelihood flows (Chen et al., 2018; Grathwohl et al., 2018): every training step solves the equation of §2, forward and backward.
- Diffusion and score-based models (Sohl-Dickstein et al., 2015; Song and Ermon, 2019; Ho et al., 2020): simulation-free, but only on paths a noising process can produce.
- Concurrent work: Liu et al. (2022, rectified flow) and Albergo and Vanden-Eijnden (2022, stochastic interpolants) arrived at similar per-example objectives independently.
Why it matters today
Three groups finding the same idea at once is a sign it was the natural next step. The three names (flow matching, rectified flow, stochastic interpolants) describe closely related recipes, and the diffusion lesson teaches them as one.
6 Experiments · original
Everyday picture
The experiments are a fair race. The same network (a U-Net) and the same settings, trained five ways: DDPM's noise loss, two versions of score matching, Flow Matching on diffusion tracks, and Flow Matching on straight tracks. If anything the race favours the older methods, which were allowed extra training iterations. Three scores are kept: how probable the model finds real test images (likelihood, in bits per dimension), how realistic its samples look (FID), and how many network calls its sampler needs (NFE). Lower is better for all three.
6.1 Likelihood and sample quality on ImageNet · original
“On both CIFAR-10 and ImageNet, FM-OT consistently obtains best results across all our quantitative measures compared to competing methods.”Lipman et al. (2022), §6.1
The results
| Training method | CIFAR-10 NLL | CIFAR-10 FID | CIFAR-10 NFE | ImageNet 64 NLL | ImageNet 64 FID | ImageNet 64 NFE |
|---|---|---|---|---|---|---|
| DDPM | 3.12 | 7.48 | 274 | 3.32 | 17.36 | 264 |
| Score matching | 3.16 | 19.94 | 242 | 3.40 | 19.74 | 441 |
| Flow Matching, diffusion paths | 3.10 | 8.06 | 183 | 3.33 | 16.88 | 187 |
| Flow Matching, optimal transport paths | 2.99 | 6.35 | 142 | 3.31 | 14.45 | 138 |
Read across the highlighted row: the straight-path model is best on every column. The NFE column is the easiest to feel: an adaptive solver, asked for the same accuracy, needs about half as many network calls on the straight-path model as on DDPM. On ImageNet 128×128 the same recipe reached an FID of 20.9, better than the unconditional GANs the paper lists (it sets aside one model, IC-GAN, that is given extra conditioning).
Bits per dimension, decoded
Everyday picture: if you used the model to compress images, how many bits would each pixel colour cost? A model that finds real images more probable compresses them into fewer bits, so lower is better.
In words: “the number of bits the model needs to describe a test image, divided by the number of colour values in it, averaged over test images.”
With the numbers: a CIFAR-10 image has d = 32 × 32 × 3 = 3,072 colour values. At 2.99 bits per dimension, one image costs 2.99 × 3,072 ≈ 9,185 bits, about 1,148 bytes, against 3,072 bytes stored raw.
In Python:
d = 32 * 32 * 3
d # → 3072
bpd = 2.99
# −log₂ p_1(x_1) = BPD × d bits per image
bits = bpd * d
round(bits) # → 9185
round(bits / 8) # → 1148
Faster training
The paper reports that Flow Matching also learns faster. For ImageNet 128, a reference diffusion model trained for 4.36 million iterations at batch size 256, while Flow Matching (with a 25% larger model) used 500,000 iterations at batch size 1.5k: about a third fewer images seen in total.
In Python:
diffusion_images = 4.36e6 * 256
fm_images = 500e3 * 1.5e3
# the share of image throughput saved
round(1 - fm_images / diffusion_images, 2) # → 0.33
Why it matters today
A method that is better on quality, likelihood, sampling cost and training cost at once is rare. That combination, with a loss simpler than diffusion's, is why straight-path training spread so quickly.
6.2 Sampling efficiency · original
“…the FM with OT model produces the best numerical error, in terms of computational cost, requiring roughly only 60% of the NFEs to reach the same error threshold as diffusion models.”Lipman et al. (2022), §6.2
Everyday picture
Following a wind with a few big steps is like driving a winding road with your eyes closed, opening them only every few seconds: on a straight road you arrive fine, on a winding one you end up in a field. Straighter paths forgive bigger steps, so they need fewer network calls for the same result.
Try it: the exact wind, a few big steps
Reading it: the four outlined squares are the data points at (±2, ±2). The dots are 300 samples, each started from the same noise every time, after N Euler steps along the exact averaged wind of equation 8 (so this is what a perfectly trained network would do, not a trained network). Faint lines show some of the journeys. With optimal transport paths and 1 step, everything lands in the middle: the averaged wind at t = 0 points at the mean of the data. By 4 steps the samples sit on the corners. Switch to diffusion paths: at 4 steps the samples are still smeared out, and it takes around 15 to get as close. The numbers are illustrative, from a four-point toy; the paper's own measurement on ImageNet 32 is the 60% quoted above.
Hover or tap the chart, or focus it and use the arrow keys, to compare the two at each step count.
Reading it: the x-axis is the number of Euler steps (network calls per sample); the y-axis is the average distance from each sample to its nearest data point, computed on the toy above. Lower is better. The solid line (optimal transport) falls steeply and reaches the floor within about 5 steps; the dashed line (diffusion paths) needs roughly 15 to 20. At a single step diffusion looks better only because its samples have barely moved from the noise, while the straight-path samples have all jumped to the centre. This is a redrawn cousin of the paper's Figure 7, where error against a 1,000-step reference is plotted against NFE on ImageNet 32.
Why it matters today
Sampling cost is paid on every image a user asks for, forever; training is paid once. That is why so much later work (rectified flow's retraining, distillation into one- or few-step students) pushes the step count down further, and why the lesson's step_sweep measures exactly this curve for a trained network.
6.3 Conditional sampling from low-resolution images · original
Everyday picture
The same recipe works when the generator is given a hint. Here the hint is a small 64×64 photo, and the task is super-resolution: produce a plausible 256×256 version.
| Model | FID ↓ | IS ↑ | PSNR ↑ | SSIM ↑ |
|---|---|---|---|---|
| Regression | 15.2 | 121.1 | 27.9 | 0.801 |
| SR3 (diffusion) | 5.2 | 180.1 | 26.4 | 0.762 |
| Flow Matching, optimal transport paths | 3.4 | 200.8 | 24.7 | 0.747 |
FID and IS judge realism, and Flow Matching wins both. PSNR and SSIM reward matching the original pixel by pixel, which is why plain regression, which blurs, tops those columns. Following the SR3 authors, the paper treats FID and IS as the better signal of quality.
Why it matters today
Conditioning slots in without changing the method: the network simply also receives the hint. Text-to-image models condition on a prompt in the same way, and guidance (see the diffusion lesson, Step 5) works on velocities just as it does on noise guesses.
7 Conclusion · original
“Furthermore, the FM framework provides an alternative view on diffusion models, and suggests forsaking the stochastic/diffusion construction in favor of more directly specifying the probability path, allowing us to, e.g., construct paths that allow faster sampling and/or improve generation.”Lipman et al. (2022), §7
Everyday picture
Diffusion said: decide how to destroy data, and learning to undo it follows. Flow Matching says: decide the road from noise to data directly, and learning to drive it follows. The first is a special case of the second.
Why it matters today
The authors expected other paths (non-round Gaussians, other kernels) to follow. The short social responsibility section adds two notes: image generators can be misused, and methods that need fewer training updates save real energy.
Appendices, briefly · original
What they contain
- A, the proofs. Theorem 1 checks the continuity equation below; Theorem 2 expands both squares and swaps the order of averaging (the one-line identity in §3.2); Theorem 3 differentiates the stretch-and-shift flow.
- C, computing probabilities. Following the flow backwards from a data point while adding up how much the wind spreads space along the way gives its log-probability, the number behind the NLL column. The spreading is estimated with random probe vectors, which keeps the estimate unbiased and cheap.
- D, diffusion's winds. Rewriting a diffusion's density equation as a continuity equation recovers the probability-flow wind, which matches Theorem 3 for both VE and VP paths.
- E, implementation. Adam, learning rates of 10−4 to 5 × 10−4, batch sizes up to 2,048, and sampling with the adaptive dopri5 solver at tolerance 10−5.
Appendix B: the continuity equation · original
Everyday picture
Watch one stretch of a busy corridor. The number of people in it goes up only if more walk in than walk out. Density times velocity is the flow of people; if that flow is larger leaving a spot than arriving, the spot empties. This bookkeeping rule is how the paper proves that a wind “generates” a crowd.
Tiny example
A steady wind of speed 2 carries the standard bell curve to the right. At x = 1, the crowd is on the down-slope, so as it moves right the density there rises at rate 0.4839; at the same time more people flow out of the spot than in, at rate −0.4839. The two cancel.
In words: “the rate at which the density at a spot changes, plus the net outflow of mass from that spot, is zero: mass is only ever moved, never made or destroyed.”
With the numbers: pt(x) = p0(x − 2t) and v = 2. At x = 1, t = 0: dp/dt = 0.4839 and div(p v) = d(2p)/dx = −0.4839; they sum to 0.
In Python:
import math
def p0(x):
return math.exp(-x * x / 2) / math.sqrt(2 * math.pi)
c, x, h = 2.0, 1.0, 1e-6
# the crowd drifting right at speed c
def p(t, x):
return p0(x - c * t)
# d/dt p_t(x), by a small finite difference
dp_dt = (p(h, x) - p(-h, x)) / (2 * h)
round(dp_dt, 4) # → 0.4839
# div(p v): in one dimension, the slope of p·c in x
div = (p(0, x + h) * c - p(0, x - h) * c) / (2 * h)
round(div, 4) # → -0.4839
abs(dp_dt + div) < 1e-6 # → True
Why it matters today
It is the one tool every proof in the paper leans on, and it is why the averaged wind of §3.1 is correct: add up the continuity equations of all the one-destination crowds and you get the continuity equation of the whole crowd.
What changed since 2022
The core, regress a velocity along a chosen path with a squared error, is unchanged. Much around it has moved:
| In the paper | Common today | Why | Lesson |
|---|---|---|---|
| Straight paths with a tiny σmin | The same with σmin = 0, as in rectified flow | One setting fewer to choose | diffusion |
| Straight conditional paths only | Retraining on the model's own (noise, sample) pairs to straighten the learned paths | Few-step and even one-step sampling | diffusion |
| A U-Net on pixels | A transformer over patches of an autoencoder's latent | Cheaper per image, and it scales like language models | diffusion |
| Adaptive dopri5 with over 100 calls | A few dozen fixed Euler steps, or a distilled few-step student | Sampling cost is paid on every image | diffusion |
| Unconditional images, super-resolution | Text-to-image and text-to-video, trained with rectified flow (Stable Diffusion 3, Esser et al. 2024) | The recipe scales and conditions easily | multimodal |
Glossary
Every term with hover guidance on this page, in one place.