An annotated companion · AI Primer

Flow Straight and Fast: Rectified Flow, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The equations are reproduced because mathematics is not copyrightable. The paper is distributed under arXiv's standard licence, which does not grant permission to republish its tables, so only a few selected rows appear here, each attributed, and every figure is redrawn from scratch and computed live in your browser. The worked numbers come from small toys built on this page, not from the paper, unless a sentence says otherwise. Read the original alongside: every section links to it.

How to read this page

Nothing on this page assumes you already know the jargon. Three things help:

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same, and each equation is followed by a table of its symbols, a sentence reading it aloud, the numbers of a tiny example, and the same example in a few lines of Python.
  • The pictures are live: press the buttons, drag the sliders, and the readouts say what changed.

Each idea is explained in the same order: an everyday picture, a tiny example you could check by hand, a diagram, then the math, and finally why it still matters.

Which way time runs. Here X0 is drawn from the starting distribution π0 (for generation, noise) and X1 from the target π1 (the data), and time runs from 0 to 1. That matches the flow matching companion and the diffusion lesson's Step 4, and is the opposite of the score SDE and DDPM companions, where time runs from data to noise.

The one-dimensional toy. A crowd of noise, π0, is a bell curve centred at 0 with spread 1; the “data”, π1, is a bell curve centred at 3 with spread 0.5. Pair them at random and the average squared distance between partners is 32 + 1 + 0.25 = 10.25. Everything about this toy can be worked out exactly, so §2 and §3 use it for their numbers. The pictures also use a two-dimensional toy of four blobs, and the noise-to-three-clusters toy of the paper's Figures 4 and 5.

Abstract

“The idea of rectified flow is to learn the ODE to follow the straight paths connecting the points drawn from π0 and π1 as much as possible.”Liu, Gong and Liu (2022), Abstract. Read the original

Everyday picture

You have to move a crowd from one side of a square to the other, and at first you pair people with destinations at random. Everyone walks in a straight line to their destination, but the lines cross everywhere. Now watch the traffic: at each spot and moment, people are moving in some average direction. Replace everyone with sleepwalkers who simply follow the traffic. The sleepwalkers never bump into each other, still arrive at the same set of destinations, and walk less in total. Pair each starting spot with where its sleepwalker ended up, and do it again: the lines barely cross any more, and a single stride gets everyone home.

What the paper claims

  • Training a flow to follow straight lines between paired points is a plain least-squares problem, with no adversarial game and no likelihood to compute.
  • The learned flow (the rectified flow) turns any pairing, or coupling, into a deterministic one that costs no more to transport, for every convex cost at once.
  • Repeating the procedure on the flow's own pairs, called reflow, provably straightens the paths, so few or even one Euler step is enough.
  • The same recipe does generation and unpaired domain transfer. On CIFAR-10 a distilled one-step model reaches FID 4.85, a record for one-step diffusion and flow models at the time.
  • The probability flow ODEs and DDIM of diffusion models are special cases with curved paths.

Why it matters today

Straight-line training is how several of today's image generators are built, Stable Diffusion 3 among them (Esser et al., 2024). The diffusion lesson builds the straight-line recipe from scratch, and the flow matching companion covers the closely related paper that appeared at the same time.

1 Introduction · original

“All these tasks can be framed unifiedly as finding a transport map between two distributions”Liu, Gong and Liu (2022), §1

Everyday picture

Generating images and translating photos into paintings look like different jobs. The paper sees one job: you have a pile of examples of one kind (noise, or photos) and a pile of another (images, or paintings), with no pairing between them, and you want a rule that turns any example of the first kind into a believable example of the second. That rule is a transport map, and the pairing it creates is a coupling. GANs learn such maps with an unstable game, likelihood models need restrictive designs, and diffusion models learn them well but need many steps to run. The paper wants the stability of diffusion and the speed of a one-step map.

Tiny example

Two starting points, −1 and +1, and two destinations, 2 and 4. Pair them in order (−1→2, +1→4) and each travels 3: an average squared distance of 9. Pair them crossed (−1→4, +1→2) and they travel 5 and 1: an average of 13. Both pairings deliver the same set of destinations; the crossed one simply wastes effort, and its lines cross. Pairing at random mixes the two and averages 11.

In words: “the transport cost of a coupling: over all pairs, the average of a cost of the jump from start to destination.” Optimal transport looks for the coupling that makes it smallest.

With the numbers: with c(x) = x2, the ordered pairing costs (32 + 32)/2 = 9 and the crossed pairing (52 + 12)/2 = 13.

In Python:

starts, ends = [-1, 1], [2, 4]
# c(x) = x²: the squared length of each jump
def c(x):
    return x ** 2
# E[c(Z_1 − Z_0)] for the ordered and the crossed pairings
ordered = sum(c(z1 - z0) for z0, z1 in zip(starts, ends)) / 2
crossed = sum(c(z1 - z0) for z0, z1 in zip(starts, ends[::-1])) / 2
ordered, crossed  # → (9.0, 13.0)

Why it matters

Framing generation as transport gives a yardstick beyond “do the samples look right”: how far, and how straight, each point travels. Straight, short journeys are exactly what make a sampler fast, which is the paper's real target.

2 Method · original

The method is short enough to state in one section; §3 then proves its properties. The paper's own order is followed here: the recipe (§2.1), what it guarantees (§2.2), and how diffusion's ODEs fit in (§2.3).

2.1 Overview · original

“By fitting the drift v with X1 − X0, the rectified flow causalizes the paths of linear interpolation Xt, yielding an ODE flow that can be simulated without seeing the future.”Liu, Gong and Liu (2022), §2.1

Everyday picture

Draw a straight line from each starting point to its partner. A point moving along its line needs to know where it is going: that is “seeing the future”, and a generator can't. So train a network to look only at where a point is and when, and to guess the line's direction. Where several lines pass through the same spot, the network can only give their average. Following those averaged arrows is an ordinary differential equation that anyone can run forwards from noise, with no knowledge of partners.

Tiny example

Noise X0 = −1 paired with data X1 = 2. A quarter of the way along, at t = 0.25, the line passes −0.25, and it moves at 2 − (−1) = 3 the whole way. The network is shown the point −0.25 and the time 0.25 and asked for 3; a guess of 2.5 costs (3 − 2.5)2 = 0.25. These are the same numbers as the diffusion lesson's flow matching step, because it is the same loss.

In words: “find the velocity field that, averaged over every pair and every moment, best matches the straight line's direction at the point where the line is at that moment.”

With the numbers: Xt = 0.25 × 2 + 0.75 × (−1) = −0.25; target 2 − (−1) = 3; a guess of 2.5 scores 0.25.

In Python:

x0, x1, t = -1.0, 2.0, 0.25
# X_t = t X_1 + (1 − t) X_0: where the line is at time t
x_t = t * x1 + (1 - t) * x0
x_t  # → -0.25
# X_1 − X_0: the line's direction, the same at every t
target = x1 - x0
target  # → 3.0
# a guess of 2.5, graded by squared error
(target - 2.5) ** 2  # → 0.25
Training (one step) Sampling, then optional reflow or distillation reflow: new pairs time t ~ U[0, 1] pairs (X₀, X₁) point Xₜ = tX₁ + (1 − t)X₀ targetX₁ − X₀ network v_θ(Xₜ, t) loss ‖X₁ − X₀ − v_θ‖² simulate dZₜ = v_θ(Zₜ, t) dtfrom Z₀ ~ π₀ to get Z₁ new pairs (Z₀, Z₁): the rectified coupling reflow: train againon the new pairs distill intoone step

Hover or tap a block. Start with pairs at the top right.

Algorithm 1 of the paper (rectified flow, with its optional reflow and distillation), drawn as a flowchart for this page.

Reading it: the top box is one training step, and it contains no simulation: a random pair and a random time give a point on the pair's line and the line's direction, the network sees only the point and the time, and the squared error is the loss. The bottom box starts once the network is trained: follow its arrows from fresh noise to get a sample, which pairs each starting point Z₀ with an end point Z₁. Those new pairs can go round again (the long wire back to the top, “reflow”), which makes the next flow straighter, or be distilled into a network that jumps from Z₀ to Z₁ in one call.

Figure 2, redrawn: crossing lines, rewired paths

starts in the upper-left blob (solid)starts in the lower-left blob (dashed)

250 random pairs between two left blobs and two right blobs; each flow is estimated from its pairs with the paper's kernel estimator (equation 5, bandwidth 0.2) and simulated live. Based on Liu, Gong and Liu (2022), Figures 2 and 3, redrawn.

Reading it: hollow circles are starting points (π0, two blobs on the left), filled dots are where they end (π1, two blobs on the right); 40 of the 250 journeys are drawn, solid if they start in the upper blob and dashed if in the lower. Random pairs: half the lines cross over to the far blob, making a big X in the middle. 1-rectified flow: the learned paths never cross; at the crossing they are “rewired”, each staying on its own side and bending to the nearer target blob. Its new pairs: join each start to where its flow ended, with straight lines: almost no crossings left. 2-rectified flow: the flow trained on those pairs is nearly straight. Now drag N down to 1. The 1-rectified flow's single step sends each point to the average end of the few training lines that start near it; with random pairs those ends are half up and half down, so the samples land in the empty gap between the two right blobs, scattered up and down and crossing each other. The 2-rectified flow's single step lands close to where its 60-step paths end. The readout gives each picture's average squared jump length (its transport cost) and its straightness (0 means perfectly straight).

Why it matters today

The loss is the same one the flow matching paper arrived at independently, with its tiny σmin set to 0. The lesson trains it in train_velocity_predictor, with flow_point as Xt and flow_target as X1 − X0, and its figure of crossing training lines and bending learned paths is the picture above.

2.2 Main results and properties · original

“It is not recommended to apply too many reflow steps as it may accumulate estimation error on vX.”Liu, Gong and Liu (2022), §2.2

The best possible arrow: an average

Everyday picture: at a busy crossroads, people heading many ways pass the same spot at the same moment. A single signpost there can only point the average way. Tiny example: in the one-dimensional toy, many lines pass through 2 at t = 0.5: from 0 to 4 (speed 4), from 1 to 3 (speed 2), from −1 to 5 (speed 6), and so on. Weighted by how likely each pairing is, their average speed is exactly 2.4.

In words: “the best velocity at a place and time is the average direction of all the lines that pass through that place at that time” (the paper's equation 2, written with the average's condition pulled out front).

With the numbers: for two independent bell curves the average is a straight-line function of x: vX(x, t) = 3 + (c/q)(x − 3t), where q = (1 − t)2 + 0.25t2 is the spread of Xt squared and c = 0.25t − (1 − t) is how X1 − X0 moves with Xt. At t = 0.5, x = 2: q = 0.3125, c = −0.375, so v = 3 − 1.2 × 0.5 = 2.4.

In Python:

t, x = 0.5, 2.0
# three of the many lines through x = 2 at t = 0.5, and their speeds
for x0, x1 in [(0, 4), (1, 3), (-1, 5)]:
    print((1 - t) * x0 + t * x1, x1 - x0)  # → 2.0 4 2.0 2 2.0 6
# the exact average for noise N(0, 1) paired at random with data N(3, 0.5²):
# q is the variance of X_t, c its covariance with X_1 − X_0
q = (1 - t) ** 2 + 0.25 * t * t
c = 0.25 * t - (1 - t)
q, c  # → (0.3125, -0.375)
3 + c / q * (x - 3 * t)  # → 2.4

Marginals are preserved, couplings get cheaper

Following the average arrows keeps the crowd at every moment exactly as the straight lines had it (Theorem 3.3, §3.1): same spread of positions at t = 0.25, 0.5, 0.9, everything. So the flow still carries π0 onto π1. And it gets there more cheaply (Theorem 3.5). The paper's picture proof, for the cost “distance travelled”, redrawn:

AAA BBB CCC DDD random pairs 4.47 + 4.47 = 8.94 rewired at P 4.47 + 4.47 = 8.94 new pairs, straight 4 + 4 = 8 A = (0, 2), B = (0, 0), C = (4, 2), D = (4, 0), P = (2, 1)

Hover or tap a part. Start with the crossing lines.

The paper's graphical proof of Theorem 3.5 for the cost c(x) = ‖x‖, redrawn with numbers for this page.

Reading it: each panel has two starting points, A and B, on its left and two destinations, C and D, on its right; the solid line is A's journey and the dashed line B's. Left: the random pairing sends A to D and B to C; the lines cross at P, each 4.47 long, 8.94 together. Middle: a flow cannot let paths cross, so at P each path turns: A ends at C and B at D. The rewired paths are the same four half-lines rearranged, so the total is still 8.94: that is the “(**)” step of the paper's proof. Right: the new pairs, A with C and B with D, joined straight. A straight line is never longer than a bent one, so these are 4 each, 8 together: the “(*)” step, the triangle inequality.

In words: “the new pairs are no farther apart, on average, than the lengths of the flow's paths, which are the old straight lines cut and re-joined, so no longer than the old pairs.”

With the numbers: old pairs √(42 + 22) = 4.472 each; rewired paths 2 × √(22 + 12) = 4.472 each; new pairs 4 each. Average: 4 ≤ 4.472 = 4.472.

In Python:

import math
A, B, C, D, P = (0, 2), (0, 0), (4, 2), (4, 0), (2, 1)
# the random pairs A→D and B→C, as straight lines
old = (math.dist(A, D) + math.dist(B, C)) / 2
# the rewired paths A→P→C and B→P→D: the same pieces
rewired = (math.dist(A, P) + math.dist(P, C) + math.dist(B, P) + math.dist(P, D)) / 2
# the new pairs A→C and B→D, straight
new = (math.dist(A, C) + math.dist(B, D)) / 2
round(new, 3), round(rewired, 3), round(old, 3)  # → (4.0, 4.472, 4.472)

Straightness, and why reflow gives it

Everyday picture: a path is straight when the direction you are moving at every moment is the same as the direction from start to finish. Tiny example: the path zt = t2 goes from 0 to 1, so the overall direction is 1, but its speed at time t is 2t: slow, then fast. Its straightness score is the average of (1 − 2t)2 over the journey: 1/3. The straight path zt = t scores 0.

In words: “straightness is the average squared gap between the velocity at each moment and the overall start-to-finish direction; zero means every path is a straight line travelled at constant speed.”

With the numbers: for zt = t2, the integral of (1 − 2t)2 from 0 to 1 is 1/3. For the one-dimensional toy's 1-rectified flow, whose paths bend (§3.4), it is 0.2146.

In Python:

n = 100_000
# the path z_t = t²: overall direction 1, velocity 2t at time t
S = sum((1 - 2 * ((i + 0.5) / n)) ** 2 for i in range(n)) / n
round(S, 4)  # → 0.3333

Theorem 3.7 promises that reflow drives straightness toward zero, at a guaranteed pace:

In words: “after K rounds of reflow, at least one of the flows has straightness no worse than the original pairs' average squared distance divided by K.” The total straightness you can ever remove is bounded by the starting cost, because each round's straightening is paid for by a drop in cost (§3.3).

With the numbers: the toy starts at 10.25, so the bound is 10.25 after one round, 5.125 after two, 1.025 after ten. It is a guarantee, and loose: the toy's first flow already has straightness 0.2146, and its second is exactly straight (§3.4).

In Python:

# E‖X_1 − X_0‖² for noise N(0, 1) paired at random with data N(3, 0.5²)
cost = 3 ** 2 + 1 + 0.5 ** 2
cost  # → 10.25
[cost / K for K in (1, 2, 10)]  # → [10.25, 5.125, 1.025]

Hover or tap the chart, or focus it and use the arrow keys, to read both curves.

Straightness and transport cost against the number of rectifications, on the two-dimensional toy of Figure 2, computed live. Based on Liu, Gong and Liu (2022), Figure 3(d), redrawn.

Reading it: the x-axis counts rectifications: 0 is the random pairing, 1 the 1-rectified flow, and so on. The y-axis is relative: the solid line is the straightness of each flow as a share of the first flow's, and the dashed line the transport cost of each coupling as a share of the random pairing's. Straightness collapses after one reflow (to a few percent of where it started) and creeps down after that; the cost drops once, when the crossing lines are removed, and then hardly moves. That matches the paper's advice that one or two reflows are enough, and that too many can pile up estimation error.

Burgers' equation. A flow is exactly straight when every point keeps its velocity for the whole trip. That makes the velocity field special: it must satisfy the inviscid Burgers' equation.

In words: “if you ride along with the flow, the velocity you feel never changes: the change at a fixed spot is exactly cancelled by the change from moving to a new spot.”

With the numbers: the field v(z, t) = z/(1 + t) carries each point along a straight line zt = z0(1 + t). At z = 2, t = 0.5: ∂tv = −0.8889 and (∂zv)v = +0.8889; they cancel.

In Python:

def v(z, t):
    return z / (1 + t)
z, t, d = 2.0, 0.5, 1e-6
# ∂_t v and ∂_z v by small finite differences
dv_dt = (v(z, t + d) - v(z, t - d)) / (2 * d)
dv_dz = (v(z + d, t) - v(z - d, t)) / (2 * d)
round(dv_dt, 4), round(dv_dz * v(z, t), 4)  # → (-0.8889, 0.8889)
abs(dv_dt + dv_dz * v(z, t)) < 1e-6  # → True

Distillation, and learning the arrows without a network

After k rounds, a network can be trained to jump from Z0 to Z1 in one call: distillation. Taking the jump as z0 + v(z0, 0) turns its loss into the rectified flow loss at t = 0 alone. The paper stresses the difference: distillation copies a coupling faithfully, while rectification changes the coupling to a cheaper, straighter one, so distillation belongs at the very end.

For low-dimensional toys the paper replaces the network with a smoothing formula that averages the lines near a point, the Nadaraya-Watson estimator (equation 5). Every picture on this page that learns a flow uses it.

In words: “to guess the arrow at z, look at every training line's position at time t; for each, the direction that would carry z to that line's end point in the time left; average those directions, weighting lines that pass close to z far more.” The kernel κh(x, z) = exp(−‖x − z‖2/2h2) sets what “close” means.

With the numbers: three training pairs (0, 4), (1, 2), (−2, 3) sit at 2, 1.5 and 0.5 at t = 0.5. At z = 1.8 with h = 0.5, their weights are 0.923, 0.835 and 0.034, and the directions to their end points are 4.4, 0.4 and 2.4. The weighted average is 2.498: mostly the two nearby lines, split nearly evenly.

In Python:

import math
pairs = [(0, 4), (1, 2), (-2, 3)]
t, z, h = 0.5, 1.8, 0.5
# κ_h(X_t, z): how close each line passes to z at time t
kappa = [math.exp(-(((1 - t) * x0 + t * x1) - z) ** 2 / (2 * h * h)) for x0, x1 in pairs]
[round(k, 3) for k in kappa]  # → [0.923, 0.835, 0.034]
# (X_1 − z)/(1 − t): the direction that reaches each end point in the time left
aims = [(x1 - z) / (1 - t) for x0, x1 in pairs]
[round(a, 1) for a in aims]  # → [4.4, 0.4, 2.4]
# ω_h = κ_h / Σ κ_h, then the weighted average
round(sum(k * a for k, a in zip(kappa, aims)) / sum(kappa), 3)  # → 2.498

Why it matters today

This section is the reason “one or two steps” is possible at all. Every sampler takes straight steps; straight paths are the only ones such steps follow without error. Reflow comes with a guarantee that learned paths get straighter, and distilling a straightened flow is a route to one-step generators with no adversarial training.

2.3 A nonlinear extension · original

“Such generalized rectified flows can still transport π0 to π1 (Theorem 3.3), but no longer guarantee to decrease convex transport costs, or have the straightening effect.”Liu, Gong and Liu (2022), §2.3

Everyday picture

Nothing forces the roads between partners to be straight. Any smooth curve from X0 to X1 works, and averaging its direction still gives a flow that moves the crowd correctly. What you lose is the two guarantees above, cheaper couplings and straightening, which relied on straight lines being the shortest roads.

Tiny example

Take the road Xt = t2X1 + (1 − t2)X0. It is the same straight line as before, but travelled slow-then-fast. With X0 = −1 and X1 = 2, at t = 0.5 it is at −0.25 (where the steady road was at t = 0.25) and moving at speed 3.

In words: “the same loss with any road: match the road's velocity at the road's position, weighting each moment by wt.” With αt = t, βt = 1 − t and wt = 1 it is equation 1 again.

With the numbers: α = 0.25, β = 0.75, α̇ = 1, β̇ = −1 at t = 0.5: Xt = 0.5 − 0.75 = −0.25, Ẋt = 2 + 1 = 3.

In Python:

t, x0, x1 = 0.5, -1.0, 2.0
# α_t = t², β_t = 1 − t², and their rates of change
alpha, beta = t * t, 1 - t * t
alpha_dot, beta_dot = 2 * t, -2 * t
alpha * x1 + beta * x0, alpha_dot * x1 + beta_dot * x0  # → (-0.25, 3.0)

Why it matters today

This one generalisation is enough to fit diffusion's samplers inside the framework, as the next subsection shows, and to compare roads on equal terms.

2.3.1 Probability flow ODEs and DDIM · original

“The linear rectified flow yields nearly straight trajectories with one step of reflow. But the trajectories of VP ODE and sub-VP ODE are curved and can not be straightened by reflowing.”Liu, Gong and Liu (2022), Figure 4 caption

Everyday picture

The score SDE paper's probability flow ODEs, and DDIM, turn out to be rectified flows along curved, unevenly paced roads. The shape of the road was never chosen for sampling: it was inherited from a noising process designed years earlier for stochastic models.

Tiny example

Halfway along the VP road (t = 0.5), the moving point is only 28% data and 96% noise (for VP the squares of the two shares add to 1, not the shares). On the straight road it is 50% of each.

In words: “the data share rises slowly and then steeply toward 1, following the exponential fade of the score SDE's VP noising read backwards; the noise share is set so the two shares' squares add to 1 (VP), or so it shrinks faster (sub-VP).” The defaults a = 19.9 and b = 0.1 are the score SDE paper's βmax − βmin and βmin.

With the numbers: at t = 0.5: α = e−1.26875 = 0.2812; VP β = 0.9597; sub-VP β = 0.9209. The straight road has α = β = 0.5.

In Python:

import math
a, b, t = 19.9, 0.1, 0.5
# α_t = exp(−¼ a (1 − t)² − ½ b (1 − t))
alpha = math.exp(-0.25 * a * (1 - t) ** 2 - 0.5 * b * (1 - t))
round(alpha, 4)  # → 0.2812
# β for VP (squares add to 1) and for sub-VP
round(math.sqrt(1 - alpha ** 2), 4), round(1 - alpha ** 2, 4)  # → (0.9597, 0.9209)

Hover or tap the chart, or focus it and use the arrow keys, to read the curves.

Reading it: the x-axis is time from noise (0) to data (1). With data share αt selected, the straight line is rectified flow, rising steadily from 0 to 1; the dashed curve is the shared α of VP and sub-VP, which stays below 0.2 until about t = 0.4 and then shoots up. With noise share βt selected, rectified flow falls steadily, while VP and sub-VP hold near 1 for the first half. Both pictures say the same thing: the diffusion roads do almost nothing early and everything late, the uneven pace the paper objects to. Based on Liu, Gong and Liu (2022), Figure 6, redrawn from equations 7 and 8.

Figures 4 and 5, redrawn: the same noise, four roads, N steps

π0 = N(0, I) and π1 a mixture of three tight bell curves (spread 0.3) on a circle of radius 3. Each road's exact averaged velocity is computed in closed form and followed with N equal Euler steps. Based on Liu, Gong and Liu (2022), Figures 4 and 5, redrawn.

Reading it: the three outlined squares are the target clusters; dots are 300 samples, all starting from the same noise, after N Euler steps; faint lines are 40 of their journeys. Rectified flow, N = 1: everything lands on the centre, the mean of the data, exactly as the paper says a single step of a 1-rectified flow must. At N = 2 the samples have already split toward the three clusters, and by 3 to 5 steps they sit on them. VP ODE: at N = 1 or 2 the samples have barely left the noise, because the road hardly moves during the early steps; it needs about 10 steps to match what rectified flow does in 3 to 5. Sub-VP is a little better than VP. VP ODE with α = t follows VP's curved roads at a steadier pace in α, as the paper proposes; on this toy that alone does not help (at 5 steps it lands further off than VP), because the roads still bend sharply near the end. The readout gives the average distance to the nearest cluster centre; about 0.37 is the floor set by the clusters' own spread.

The VE ODE

The variance exploding ODE uses αt = 1 and a noise share that shrinks from σmax to almost 0, so its starting crowd is a very wide bell curve. Its roads are straight (the direction ξ never changes) but travelled at a very uneven pace; the paper leaves it out of its toys because it cannot start from the standard bell curve.

Why it matters today

The paper's conclusion is a design rule that stuck: the ODE can be learned directly, any road can be chosen, the starting distribution is free, and the canonical choice is the straight line at constant speed. Its score SDE ancestor derived its ODE from an SDE; this paper, and flow matching, choose the road first.

3 Theoretical analysis · original

Every result below is illustrated on the one-dimensional toy: noise N(0, 1) paired at random with data N(3, 0.52). For this toy the rectified flow can be solved exactly: a point starting at z0 is at Zt = 3t + z0√q(t), with q(t) = (1 − t)2 + 0.25t2, and so it ends at Z1 = 3 + 0.5z0.

3.1 The marginal preserving property · original

“Because Zt is driven by the same velocity field vX, its marginal law Law(Zt) solves the very same equation with the same initial condition (Z0 = X0).”Liu, Gong and Liu (2022), §3.1

Everyday picture

Film a crowd crossing a square from above and blur the individuals: you see a smooth density moving. That density only changes as fast as people flow in and out of each spot, and the flow at each spot is the density times the average velocity there (the continuity equation). The straight-line walkers and the sleepwalkers have the same average velocity everywhere, by construction, and start as the same crowd. So the film looks the same, frame by frame, even though individuals take different routes.

Tiny example

At t = 0.5, the straight lines' crowd X0.5 = 0.5X0 + 0.5X1 has mean 1.5 and variance 0.25 + 0.0625 = 0.3125. Push 10,000 noise samples through the rectified flow to t = 0.5 with small Euler steps: their mean is 1.5 and their variance 0.31.

In words: “at every moment, the flow's crowd is distributed exactly as the straight lines' crowd”; at t = 1 that means the flow delivers exactly the data distribution (Theorem 3.3).

With the numbers: mean 1.5 and variance 0.3125 for the lines; 1.5 and 0.31 for 10,000 simulated flow paths (the small gap is sampling noise and step error).

In Python:

import random
random.seed(0)
# q(t) = Var X_t and c(t) = Cov(X_1 − X_0, X_t) for the toy
def q(t):
    return (1 - t) ** 2 + 0.25 * t * t
def c(t):
    return 0.25 * t - (1 - t)
zs = [random.gauss(0, 1) for _ in range(10_000)]
steps = 200
h = 0.5 / steps
# follow v^X(z, t) = 3 + (c/q)(z − 3t) from t = 0 to 0.5
for i in range(steps):
    t = i * h
    zs = [z + h * (3 + c(t) / q(t) * (z - 3 * t)) for z in zs]
mean = sum(zs) / len(zs)
var = sum((z - mean) ** 2 for z in zs) / len(zs)
round(mean, 2), round(var, 2), q(0.5)  # → (1.5, 0.31, 0.3125)

Why it matters today

This is the property that makes any of it a generator: the flow is guaranteed to deliver π1 if the averaged velocity is learned well. It holds for every road, straight or not, which is why diffusion's ODEs work too.

3.2 Reducing convex transport costs · original

“The proof is based on elementary applications of Jensen's inequality.”Liu, Gong and Liu (2022), §3.2

Everyday picture

A convex cost punishes long jumps at least in proportion: squared distance, plain distance, distance to the fourth power. The proof uses one fact twice, Jensen's inequality: averaging first and then paying the cost is never dearer than paying the cost first and then averaging. The flow averages directions, so its pairs can only be cheaper.

Tiny example

In the toy, random pairs cost 10.25 in squared distance; the rectified pairs (z0, 3 + 0.5z0) cost 32 + 0.25 = 9.25. In plain distance, 3.0025 against 3.0000: smaller again, though barely, because both jumps are mostly the same 3 to the right.

In words: “for any convex cost at all, the rectified pairs cost no more than the pairs you started with” (Theorem 3.5). No particular cost is being minimised: all of them go down together.

With the numbers: X1 − X0 is a bell curve with mean 3 and variance 1.25; Z1 − Z0 = 3 − 0.5z0 has mean 3 and variance 0.25. Squared: 10.25 ≥ 9.25. Absolute: 3.0025 ≥ 3.0000.

In Python:

import math
# E[c] for a bell curve of mean mu and spread s, for c(x) = x² and c(x) = |x|
def e_square(mu, s):
    return mu ** 2 + s ** 2
def e_abs(mu, s):
    tail = 0.5 * (1 + math.erf(-mu / s / math.sqrt(2)))
    return s * math.sqrt(2 / math.pi) * math.exp(-mu ** 2 / (2 * s ** 2)) + mu * (1 - 2 * tail)
# X_1 − X_0 ~ N(3, 1.25) for random pairs; Z_1 − Z_0 = 3 − 0.5 z_0 ~ N(3, 0.25) after rectifying
old, new = (3, math.sqrt(1.25)), (3, 0.5)
round(e_square(*new), 4), round(e_square(*old), 4)  # → (9.25, 10.25)
round(e_abs(*new), 4), round(e_abs(*old), 4)  # → (3.0, 3.0025)

Why it matters today

Cheaper couplings mean shorter journeys, and deterministic ones mean each noise sample has one definite image: a usable latent space for editing, which the paper uses in §5.2.

3.3 The straightening effect · original

“Then, we prove that recursive rectification straightens the coupling and its related flow with a O(1/k) rate, where k is the number of rectification steps.”Liu, Gong and Liu (2022), §3.3

Everyday picture

A coupling is straight when rectifying it changes nothing: its straight lines never cross, so the average arrow at every spot is the one line through it, and the flow simply follows the lines. Each round of reflow removes crossings, and every bit of crossing it removes is paid for from a fixed budget, the starting cost. A budget that only goes down cannot keep paying forever, so the crossings, and the bending, must fade.

Tiny example

In the toy, one rectification lowers the squared cost from 10.25 to 9.25, a drop of exactly 1. That drop splits into two parts: how much the new flow bends, S = 1 − π/4 = 0.2146, and how much the old straight lines crossed, V = π/4 = 0.7854.

In words: “V measures how much the straight lines disagree with their own average where they meet, which is zero exactly when no two lines cross; and one round of rectification lowers the squared cost by the new flow's bending plus the old lines' crossing.” Summing this over rounds gives Theorem 3.7's bound.

With the numbers: V integrates 1.25 − c2/q; S integrates (0.5 + c/√q)2. They come to 0.7854 and 0.2146, and add to 1.0 = 10.25 − 9.25.

In Python:

import math
def q(t):
    return (1 - t) ** 2 + 0.25 * t * t
def c(t):
    return 0.25 * t - (1 - t)
n = 100_000
ts = [(i + 0.5) / n for i in range(n)]
# V: variance of X_1 − X_0 (1.25) minus the part the average arrow explains, c²/q
V = sum(1.25 - c(t) ** 2 / q(t) for t in ts) / n
# S: the flow's velocity c/√q · z_0 + 3 against its overall direction 3 − 0.5 z_0
S = sum((0.5 + c(t) / math.sqrt(q(t))) ** 2 for t in ts) / n
round(V, 4), round(S, 4), round(V + S, 4)  # → (0.7854, 0.2146, 1.0)
round(math.pi / 4, 4)  # → 0.7854

A note on the source: equation 13 in the paper prints the cost drop without the squares, as E‖X1 − X0‖ − E‖Z1 − Z0‖; the sentence before it (take c(x) = ‖x‖2) and the equation after it both have them, and only the squared version is true. The toy confirms it: the plain-distance drop is 0.0025, nowhere near 1.

Reading it: time runs left to right; height is position. Nine noise samples, evenly spread from −2 to 2, start on the left; the data crowd is centred at 3 with spread 0.5. Random pairs: each start is joined by a straight line to a random data point; the lines criss-cross. 1-rectified flow: the exact flow from the formula above. The paths never cross; they first squeeze together (the crowd narrows from spread 1 toward 0.45 around t = 0.8) and then fan out again to spread 0.5, so they bend. The order of the starts is kept: the lowest noise goes to the lowest data. 2-rectified flow: the same starts and the same ends, joined straight, because in one dimension the ordered pairing's lines never cross: one reflow makes the flow exactly straight.

Why it matters today

The straightening theorem is why reflow works in practice: the paper's CIFAR-10 flows go from recognisable-but-blurry images at one step to clear ones after a single reflow (§5.2).

3.4 Straight vs. optimal couplings · original

“Hence, it is “easier” to find a straight coupling than a c-optimal couplings.”Liu, Gong and Liu (2022), §3.4

Everyday picture

Straight and optimal are different goals. Optimal means the cheapest pairing for one chosen cost. Straight only means no two lines cross. Every optimal pairing is straight (if two lines crossed, swapping partners would be cheaper), but not every straight pairing is optimal, except on a line.

Tiny example

On a line (one dimension), a pairing is straight exactly when it keeps the order: the lowest start goes to the lowest destination, and so on. That ordered pairing is the cheapest for every convex cost at once. The toy's rectified pairing, z0 → 3 + 0.5z0, keeps the order, so it is both straight and optimal: 9.25 is the least squared cost any pairing of these two crowds can reach. In two or more dimensions, rectification finds a straight pairing, which need not be the cheapest; the paper points to follow-up work that restricts the velocity to the gradient of a function to reach the squared-cost optimum.

Why it matters today

For fast sampling, any straight pairing is as good as any other: all of them can be followed exactly in one step. The paper argues that chasing exact optimal transport is therefore unnecessary for generation, and also shows that an earlier conjecture, that DDIM's pairing is optimal for the squared cost, was already known to be false.

3.5 Denoising diffusion models and probability flow ODEs · original

“We prove that the probability flow ODEs (PF-ODEs) of [73] can be viewed as nonlinear rectified flows in (6) with Xt = αtX1 + βtξ.”Liu, Gong and Liu (2022), §3.5

Everyday picture

A diffusion model's ODE is trained with a target written in the language of noise processes: a shrink rate ηt and a noise strength σt. Proposition 3.11 translates that target into road language, and finds it is exactly the velocity of the road Xt = αtX1 + βtξ: the diffusion ODE was a rectified flow all along.

Tiny example

On the VP road at t = 0.5, with data X1 = 1 and noise ξ = 0.5: the road's own velocity is α̇X1 + β̇ξ = 1.4129 − 0.2070 = 1.2059. The diffusion ODE's training target, built from η0.5 = −5.025 and σ0.52 = 10.05, is also 1.2059. (That σ2 is the score SDE paper's β(0.5), read in reversed time.)

In words: “the probability flow's training target, a shrink term plus a noise term, equals the road's velocity once the shrink rate and noise strength are written in terms of the road's two shares.”

With the numbers: α = 0.2812, α̇ = 1.4129, β = 0.9597, β̇ = −0.4140; η = −5.025, σ2 = 10.05; Xt = 0.2812 + 0.4798 = 0.7610; Ỹ = 5.025 × 0.7610 − (10.05/1.9193) × 0.5 = 1.2059, the same as α̇ × 1 + β̇ × 0.5.

In Python:

import math
a, b, t = 19.9, 0.1, 0.5
u = 1 - t
# the VP road: α, β and their rates of change
alpha = math.exp(-0.25 * a * u * u - 0.5 * b * u)
alpha_dot = alpha * (0.5 * a * u + 0.5 * b)
beta = math.sqrt(1 - alpha ** 2)
beta_dot = -alpha * alpha_dot / beta
# η and σ² written in terms of the road
eta = -alpha_dot / alpha
sigma2 = 2 * beta ** 2 * (alpha_dot / alpha - beta_dot / beta)
round(eta, 3), round(sigma2, 2)  # → (-5.025, 10.05)
x1, xi = 1.0, 0.5
x_t = alpha * x1 + beta * xi
# the diffusion ODE's target and the road's velocity
y_tilde = -eta * x_t - sigma2 / (2 * beta) * xi
round(y_tilde, 4), round(alpha_dot * x1 + beta_dot * xi, 4)  # → (1.2059, 1.2059)

Two notes on the source. The proof's second line writes the shrink rate with a dot (η̇t) where the rest of the section, and the computation above, use ηt itself; and the nonlinear rectified flow's definition in §2.3 prints E[Ẋt | Xt = t] where the condition must be Xt = z. Both read as typos; neither changes a result.

Why it matters today

Seen this way, the difference between a diffusion ODE and a rectified flow is only the road: curved and unevenly paced, or straight and steady. That is the whole case for training on straight roads, and it is why the DDIM companion's sampler and the lesson's euler_step are the same kind of step on different roads.

4 Related works and discussion · original

“The success of the denoising diffusion models may be mainly attributed to the simple and stable optimization-based training procedure that allows us to avoid the instability issues and the need of case-by-case tuning of GANs, rather than the presence of diffusion noises.”Liu, Gong and Liu (2022), §4

Everyday picture

The discussion places rectified flow among the ways to learn a transport map, and then makes a provocative argument: the noise in diffusion models was a means, not the point. What made them work was a stable regression loss, and an ODE can have that without the noise.

The main comparisons

  • One-step models. GANs are unstable and can mode-collapse; VAEs and normalizing flows need approximations or constrained designs. Reflow plus distillation is offered as a fourth route.
  • Training ODEs by likelihood (neural ODEs) needs the ODE simulated in every training step and back-propagation through all of it, and is under-specified: infinitely many ODEs deliver the same final distribution. Fixing the road in advance, as rectified flow does, removes both problems.
  • Probability flow ODEs and DDIM avoid those costs but inherit roads from SDE theory that are “unnecessarily restrictive and complicated”.
  • ODEs versus SDEs. ODEs are simpler, faster to simulate, as easy to run backwards as forwards, give deterministic pairs (a useful latent space), are no harder to train, and represent the same crowds; SDEs may still suit truly noisy data or time-correlation structure.
  • Optimal versus straight transport. For fast inference any straight coupling is as good as the optimal one.

Why it matters today

Whether or not the noise was ever essential, this is the framing the flow matching companion and the diffusion lesson use: an ODE trained by regression, with diffusion's roads as one choice among several. The lesson teaches DDPM, DDIM and flow matching side by side as one family.

5 Experiments · original

Every experiment follows Algorithm 1: train on independent pairs, simulate the flow to get new pairs, reflow, and optionally distill. Flows are simulated with plain Euler steps of size 1/N, or with an adaptive solver (RK45) for full-quality numbers.

5.1 Toy examples · original

Everyday picture

To see the theory without a neural network's quirks, the toys use the kernel estimator of equation 5, averaged over the 100 nearest training lines with bandwidth 1 (the paper finds results insensitive to both). A small network also works (two hidden layers of 64), fitting the averaged arrows a little less crisply; the paper finds that stronger L2 regularization, which makes the network smoother, helps straighten the flow further.

How straightness and cost were measured

Straightness is equation 3 estimated from simulated paths, as in Figure 2 above. For cost, the paper compares each coupling with the best possible pairing of the same end points, found by solving the discrete optimal transport problem, and notes a trap: in high dimensions that relative cost is near zero even for a random network, which is how an earlier study came to believe DDIM was optimal.

Why it matters today

Low-dimensional pictures like these are still the fastest way to see what a sampler does to paths; the redrawn figures on this page are built the same way.

5.2 Unconditioned image generation · original

“In particular, the distilled 2-rectified flow achieves an FID of 4.85, beating the best known one-step generative model with U-net architecture, 8.91 (TDPM, Table 1 (b)).”Liu, Gong and Liu (2022), §5.2

Everyday picture

On CIFAR-10 the paper uses the score SDE paper's DDPM++ U-Net and code, so the comparison with the VP and sub-VP ODEs is like for like: same network, different road. Quality is measured by FID and Inception score, and diversity by recall.

Selected rows of Table 1(a), reproduced with attribution (Liu, Gong and Liu, 2022): CIFAR-10, all with the DDPM++ network. NFE is the number of network calls per image. In the one-step rows, the number in brackets is after distillation
MethodNFEFID ↓Recall ↑
1-rectified flow, one Euler step1378 (6.18)0.0 (0.45)
2-rectified flow, one Euler step112.21 (4.85)0.34 (0.50)
3-rectified flow, one Euler step18.15 (5.21)0.41 (0.51)
VP ODE, one Euler step1451 (16.23)0.0 (0.29)
1-rectified flow, adaptive solver1272.580.57
VP ODE, adaptive solver1403.930.51
VP SDE, 2,000 Euler steps20002.550.58

Read the one-step rows first: without reflow, one step is useless (FID 378, the blurred average image of Figure 5's first panel); after one reflow it is 12.21 without any distillation, and 4.85 with it. With a full solver, the 1-rectified flow is the best ODE on the page (FID 2.58 with 127 calls) and close to the 2,000-step SDE. The paper also reports the trade-off honestly: reflow helps a lot below about 80 steps but slightly hurts with many steps, because each round adds some estimation error.

Seeing straightness in pictures

Figure 10 of the paper extrapolates from the middle of a path to its end with one straight jump. On a straight path that guess never changes as you move along it.

In words: “from where the path is now, jump straight to the end, at the current velocity, for all the time that is left.”

With the numbers: in the one-dimensional toy, the 1-rectified flow from z0 = 1 is at 2.0590 at t = 0.5 with velocity 2.3292, so the guess is 3.2236, though the path really ends at 3.5. On the 2-rectified flow the path is the straight line from 1 to 3.5, and the guess is 3.5 at every t.

In Python:

import math
def q(t):
    return (1 - t) ** 2 + 0.25 * t * t
def c(t):
    return 0.25 * t - (1 - t)
t, z0 = 0.5, 1.0
# the 1-rectified flow: z_t = 3t + z_0 √q(t), velocity 3 + (c/q)(z_t − 3t)
z_t = 3 * t + z0 * math.sqrt(q(t))
v = 3 + c(t) / q(t) * (z_t - 3 * t)
round(z_t, 4), round(v, 4)  # → (2.059, 2.3292)
round(z_t + (1 - t) * v, 4), 3 + 0.5 * z0  # → (3.2236, 3.5)
# the 2-rectified flow is the straight line from z_0 to 3.5
z2 = (1 - t) * z0 + t * 3.5
z2 + (1 - t) * (3.5 - z0)  # → 3.5

On cat faces, the 2-rectified flow's extrapolated image hardly changes along the path, and even the 1-rectified flow's is recognisable by t ≈ 0.1, where the sub-VP ODE's needs t ≈ 0.6. The paper also shows 256 × 256 images (bedrooms, churches, faces, cats) from the 1-rectified flow, and simple edits: run the flow backwards from a stitched-together image to its noise, nudge the noise toward more probable values, and run forwards again.

Why it matters today

This table is the evidence behind “reflow, then distill” as a route to few-step generation, and the backwards-then-forwards edit shows the practical value of a deterministic, low-cost coupling.

5.3 Image-to-image translation · original

Everyday picture

Set π0 to photos of cats and π1 to photos of wild animals, with no pairs, and the same algorithm learns to turn one into the other. There is no need for the cycle-consistency penalty of CycleGAN, because an ODE run backwards returns exactly where it started. But the goal here is not an exact match to π1: it is to change the style while keeping the subject, so the loss only counts errors that change the style.

Tiny example

A two-number “image” whose first number is its style and second its content. A pair moves from (0, 0) to (2, 1); the network guesses the velocity (1.5, 3.0). The plain loss counts both errors: 0.52 + 22 = 4.25. With the style feature h(x) = first number, whose slope is (1, 0), only the style error counts: 0.52 = 0.25.

In words: “the rectified flow loss, but measured only in the directions that change the style features, weighted by how much they change them.” The paper takes h from a classifier trained to tell the two domains apart.

With the numbers: the error is (0.5, −2.0); plain loss 4.25; style-weighted loss (1 × 0.5 + 0 × (−2.0))2 = 0.25.

In Python:

x0, x1, v = (0.0, 0.0), (2.0, 1.0), (1.5, 3.0)
# ∇h: the style feature is the first number, so its slope is (1, 0)
grad_h = (1.0, 0.0)
err = [x1[i] - x0[i] - v[i] for i in range(2)]
err  # → [0.5, -2.0]
# the plain loss, and the loss seen through ∇h
sum(e * e for e in err), sum(g * e for g, e in zip(grad_h, err)) ** 2  # → (4.25, 0.25)

Why it matters today

Generation and translation become the same algorithm with a different π0, which is the paper's unifying claim made concrete. The results (cats to wild animals, faces to painted portraits and back) are shown as images, one step being enough after one reflow.

5.4 Domain adaptation · original

Everyday picture

A classifier trained on photos stumbles on sketches. Instead of retraining it, move the sketches' features into the photos' feature distribution with a rectified flow, and classify the moved features. Here π0 and π1 are the test and training features from a pretrained model's last hidden layer, and the flow is run with 100 Euler steps.

Selected columns of Table 2, reproduced with attribution (Liu, Gong and Liu, 2022): accuracy (%) of the transferred test data; higher is better. ERM is plain training; CORAL is the strongest baseline listed
DatasetERMCORAL1-rectified flow
OfficeHome66.5 ± 0.368.7 ± 0.369.2 ± 0.5
DomainNet40.9 ± 0.141.5 ± 0.241.4 ± 0.1

The paper's own wording is careful: “better or on par with” the best baseline. On DomainNet the difference is inside the error bars.

Why it matters today

It shows that the recipe is not tied to images or to noise: any two piles of vectors can be transported, with no pairs and no adversary.

Appendix, briefly · original

What it contains

  • CIFAR-10. Adam at learning rate 2 × 10−4, dropout 0.15, and an exponential moving average of the weights with rate 0.999999. For each reflow, 4 million (z0, z1) pairs are generated and the previous model is fine-tuned on them for 300,000 steps.
  • Few-step distillation. To make a k-step generator, training times are drawn only from {0, 1/k, …, (k − 1)/k}; for one step, the squared error is replaced by LPIPS, which worked better.
  • Translation and adaptation. AdamW with weight decay 0.1, 512 × 512 images, 80% of each dataset for training; extra figures of latent interpolation, editing and reconstruction.

Why it matters

The reflow data cost is real: millions of full simulations of the previous flow before retraining. That price is paid once, at training time, in exchange for cheap sampling forever after.

What changed since 2022

The straight-line loss became a standard; the surrounding recipe kept moving.

In the paperCommon todayWhyLesson
A U-Net on 32 × 32 pixelsA transformer over patches of an autoencoder's latent, trained with the straight-line loss (Stable Diffusion 3, Esser et al., 2024)The recipe scales and conditions on text easilydiffusion
Reflow on millions of simulated pairsReflow and distillation among several ways to reach few-step samplingSampling cost is paid on every imagediffusion
Rectified flow and flow matching as separate papersTaught as one method: regress the velocity of a straight roadThe losses are the samediffusion

Glossary

Every term with hover guidance on this page, in one place.