An annotated companion · AI Primer

Generative Adversarial Nets, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. Equations are reproduced with every symbol decoded. The paper's figures are not reproduced: Figure 1 and Table 2 are redrawn from scratch as interactive pictures, the sample images of Figures 2 and 3 are described in words, and three rows of Table 1 are quoted with attribution. Worked numbers are our own, and made-up ones are labelled illustrative. Read the original alongside: every section heading links to it.

How to read this page

Nothing here assumes you already know the jargon.

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same, and every equation is followed by a table of its symbols, a sentence reading it aloud, the numbers from a tiny example, and the same numbers in plain Python.
  • The pictures are live. The redrawn Figure 1 in §3 is the fastest way to see the whole idea: drag the generator and watch the detective react.

Each idea climbs the same ladder: an everyday picture, a tiny example you can check by hand, a diagram, the math, and why it matters today. The GAN lesson builds both networks from scratch in NumPy, trains them on a ring of eight little clouds, and shows every way the game goes wrong.

Abstract · original

“The training procedure for G is to maximize the probability of D making a mistake.”Goodfellow et al. (2014), Abstract

Everyday picture

A forger prints banknotes and a detective inspects them. Every time the detective catches a fake, the forger learns what gave it away. Every time the forger improves, the detective has to look harder. Nobody ever writes down what a real banknote looks like, yet after enough rounds the forger prints notes the detective cannot tell from the real thing. The forger is the generator G, the detective is the discriminator D, and the pair is a generative adversarial network.

What the paper claims

  • A new way to train a generative model: two networks play a minimax game, one trying to make a single number big and the other trying to make it small.
  • If both players could be any function at all, the game has exactly one solution: G reproduces the data's distribution and D says ½ everywhere.
  • When both players are multilayer perceptrons, everything trains with plain backpropagation. No Markov chains and no approximate inference are needed, either to train or to draw samples.

Why it matters today

For several years, GANs made the sharpest synthetic images anyone had seen. Diffusion models have since taken over most image generation, but the adversarial idea lives on inside image and audio decoders and in fast distilled samplers, and the instabilities this paper first met show up wherever two models are trained against each other.

1 Introduction · original

“The generative model can be thought of as analogous to a team of counterfeiters, trying to produce fake currency and use it without detection, while the discriminative model is analogous to the police, trying to detect the counterfeit currency.”Goodfellow et al. (2014), §1

Everyday picture

Telling cats from dogs is a sorting job: look at a picture, give a label. By 2014 deep networks were superb at sorting jobs, which the paper calls discriminative models. Drawing a new cat is much harder, because the model must know what every pixel of a plausible cat looks like at once. The older ways of doing it needed sums over every possible image, which nobody can compute. The paper's trick is to turn drawing into sorting: train a sorter (the detective) and use its verdicts to teach a drawer (the forger).

Tiny example

The detective only ever has to answer one yes-or-no question: is this real? That is the same job as a spam filter. If it says 0.9 for a real photo and 0.2 for a fake, it is doing well. The forger never sees a real photo at all. All it gets is the detective's answer on its own fakes, and the direction that would have made that answer higher.

Why it matters

The paper's bet is that the machinery that made classifiers work, backpropagation, dropout and piecewise linear units, can be reused unchanged to build a generator. That bet is why GANs scaled so quickly: anything you could train as a classifier, you could now turn into a detective. The GAN lesson's first section builds both players as tiny networks.

2 Related work · original

“Compared to GSNs, the adversarial nets framework does not require a Markov chain for sampling.”Goodfellow et al. (2014), §2

Everyday picture

Imagine judging how likely one particular photo is by comparing it against every photo that could ever exist. Many generative models of the time, such as Boltzmann machines, worked like that: they give each image a score, and to turn scores into probabilities they must divide by the total score of all images, the partition function. That total is impossible to compute exactly, so it was estimated with long random walks called Markov chain Monte Carlo, which get stuck when the data has several separate peaks.

Tiny example

A tiny black-and-white image of 28 × 28 pixels has 784 pixels, each on or off, so there are 2784 possible images. That number has 237 digits. A partition function over such images is a sum of 2784 terms. A GAN never needs it: its generator makes a sample in one forward pass, and its detective only compares.

In Python:

# every 28 × 28 black-and-white image: 2 choices for each of 784 pixels
images = 2 ** (28 * 28)
# how many digits that count has
len(str(images))  # → 237
training sampling p(x) Directed models Undirected models Generative autoencoders Adversarial nets inference easy estimate MCMC chain estimate trade-off chain Parzen sync D, G easy Parzen green: no difficulty · amber: hard · yellow: only estimated

Hover or tap a cell. Start with the bottom row, the paper's own method, and compare it with the rows above.

Table 2 of the paper, redrawn and summarised in our own words: the main difficulties of four families of deep generative models. Based on Goodfellow et al. (2014), Table 2.

Reading it: each row is a family of generative models and each column one job you might want from it: training it, drawing a new sample from it, and computing p(x), the probability it gives a particular image. Green means no difficulty, amber a real obstacle, yellow a number that can only be estimated. Read the bottom row last. Adversarial nets turn the sampling column green (one forward pass, no chain), like directed models, and they escape the chains and inference that make the upper rows hard to train. The price is a new amber cell: the two networks must be kept in step. Nobody can compute p(x) for the top rows either, but the last two rows have no formula for it at all, so it has to be estimated from samples (the Parzen method of §5).

Two close relatives

  • Noise-contrastive estimation also fits a model by telling real data from fakes. But its fakes come from fixed noise that never improves, so once the model is roughly right the job becomes too easy and learning slows. In a GAN the fakes come from the generator, so they keep getting harder to spot.
  • Variational autoencoders, published the same year, also train a generator by backpropagating into it. They need a second network to do inference; a GAN needs a detective instead. The autoencoders lesson builds one.

Why it matters today

The two families that survived are the two that avoid Markov chains during sampling: GANs and VAEs, and later diffusion models, which do take many steps but each one is a plain forward pass through a network rather than a random walk that must mix.

3 Adversarial nets · original

Everyday picture

The forger starts from a handful of dice rolls, the noise z, and turns them into a banknote G(z). Different rolls, different note. The detective takes a note and outputs one number, D(x): how likely it thinks the note came from the real pile rather than from the forger.

The game (equation 1) · original

Everyday picture

Picture a scoreboard. The detective earns points for confident correct calls and loses a lot for confident mistakes. The detective wants the score as high as possible; the forger wants it as low as possible. One number, two players pulling it opposite ways.

Tiny example

The detective sees two real notes and two fakes. Scores use the logarithm: log 1 = 0, and the log of a small number is very negative, so a confident mistake costs far more than a hesitant one.

NoteReal?Verdict DWhat is scoredValue
1real0.9log D = log 0.9−0.105
2real0.8log D = log 0.8−0.223
3fake0.2log(1 − D) = log 0.8−0.223
4fake0.4log(1 − D) = log 0.6−0.511

The real notes average −0.164, the fakes −0.367, total −0.53. A perfect detective would score 0. A detective reduced to a coin flip, ½ on everything, scores log ½ + log ½ = −1.386.

The math

In words: “average the log of the detective's verdicts on real data, add the average log of one minus its verdicts on fakes; the detective picks its weights to push this up, and the forger picks its weights to push the detective's best score down.”

With the numbers: (log 0.9 + log 0.8) / 2 + (log 0.8 + log 0.6) / 2 = −0.164 + (−0.367) = −0.53. The expectation 𝔼 is just “the average over many draws”, here over two.

In Python:

import math
# D(x) on two real notes, D(G(z)) on two fakes
D_x, D_Gz = [0.9, 0.8], [0.2, 0.4]
# E over real x of log D(x)
real_term = sum(math.log(d) for d in D_x) / len(D_x)
round(real_term, 3)  # → -0.164
# E over noise z of log(1 − D(G(z)))
fake_term = sum(math.log(1 - d) for d in D_Gz) / len(D_Gz)
round(fake_term, 3)  # → -0.367
# V(D, G)
round(real_term + fake_term, 2)  # → -0.53
# a coin-flip detective: D = 1/2 on everything
round(math.log(0.5) + math.log(0.5), 3)  # → -1.386

Look closely and −V is the binary cross-entropy of an ordinary classifier that labels real notes 1 and fakes 0. The detective is a plain classifier. What is new is that its second class keeps changing underneath it.

Why it matters today

Equation 1 is the whole method in one line, and it is still the starting point of every GAN loss. The value_function in the lesson computes it from a batch of verdicts.

Figure 1 in one dimension · original

Everyday picture

Lay every possible banknote along a ruler. The real notes pile up in one bump on the ruler; the forger's notes pile up in another. The detective's verdict is a curve over the ruler: high where real notes dominate, low where fakes dominate. The forger then slides its pile towards where the curve is high.

Tiny example

The forger's recipe here is G(z) = μ + σ · (the bell-curve value at position z), with z drawn evenly between 0 and 1, as in the paper's figure. With μ = 0.2 and σ = 0.9, the middle roll z = 0.5 lands on x = 0.2, and rolls near the ends of the z line land far out in the tails. Evenly spaced rolls land close together near the middle of the pile and spread out at its edges: the generator contracts where pg is dense and expands where it is sparse, exactly as the paper's caption says.

Stage:
0.20 0.90
real data pdata (dotted)generator pg (solid)detective D(x) (dashed)

Pick a stage, drag the sliders, or hover the picture to read the three curves at any x.

Reading it: the bottom line is the noise z, from 0 on the left to 1 on the right, with eleven evenly spaced rolls. Each arrow carries one roll up to the point x = G(z) on the data line above it. The dotted bump is the real data, the solid bump is where the generator's samples pile up, and the dashed curve is the detective's verdict, on a scale from 0 (bottom) to 1 (top). Step through the stages. In (a) the forger is close and the detective only partly right. In (b) the detective has been trained to its best, which is high over the real bump and low where only fakes live. In (c) the forger has taken a step towards where that verdict was high, so the old dashed curve is now out of date. In (d) the bumps coincide and the verdict is flat at ½: the detective can only guess. Drag μ and σ to put the forger anywhere; the dashed curve then shows the best possible detective for that forger. The densities are drawn to their own scale, and the stage settings are illustrative, chosen to match the paper's four panels.

Why it matters today

This picture is the paper's argument in miniature: the detective's curve tells the forger which way to move its mass, and the game stops only when there is nothing left to tell. The GAN lesson draws the same idea in two dimensions, with a trained detective's verdict as a coloured background behind the forger's samples.

Algorithm 1: taking turns · original

“The number of steps to apply to the discriminator, k, is a hyperparameter. We used k = 1, the least expensive option, in our experiments.”Goodfellow et al. (2014), Algorithm 1

Everyday picture

Nobody can solve “min over G of max over D” in one go. So the players take turns: the detective studies for a little while, then the forger practises once against the detective as it now is, and round it goes. Training the detective to perfection before each forger move would be too slow, and on a finite pile of real notes the detective would just memorise them (overfit).

Tiny example

Take the four notes from the table above as one round with a minibatch of m = 2. The detective's objective is the same −0.53. Nudging each of its scores (the number before the sigmoid) changes the objective at these rates: +0.05 and +0.1 on the two real notes (so it raises them), −0.1 and −0.2 on the fakes (so it lowers them). Backpropagation passes those rates on to every weight θd. Then the forger takes its turn on two fresh rolls of noise.

for each training iteration k times (k = 1 in the paper) sample m noise zmake m fakes sample m real xfrom the data update D: climbits stochastic gradient sample m fresh z update G: descendits stochastic gradient

Hover or tap a step. Start at the top left with the noise.

Algorithm 1 of the paper, redrawn as a flow: minibatch stochastic gradient training of generative adversarial nets. Based on Goodfellow et al. (2014), Algorithm 1.

Reading it: the outer dashed box is one training iteration, and the inner one is the detective's turn, repeated k times. In the detective's turn a fresh batch of noise becomes fakes, a batch of real examples arrives, and both feed one step of gradient ascent for D (blue). Then, outside the inner box, the forger draws fresh noise and takes one step of gradient descent (green) against the detective as it is after its step. The long arrow on the right starts the next iteration. With k = 1, as in the paper's experiments, the two players simply alternate.

The math

The detective's step: climb this, with respect to its own weights θd.

In words: “over a minibatch of m real examples and m fakes, average the log of the verdicts on the reals plus the log of one minus the verdicts on the fakes, and find which way each of the detective's weights should move to raise that average.”

With the numbers: the average is −0.53, as above. Written as D = σ(a), the log of a real verdict rises at rate 1 − D per unit of score and the log of one minus a fake verdict at rate −D, each divided by m = 2: (1 − 0.9)/2 = 0.05, (1 − 0.8)/2 = 0.1, −0.2/2 = −0.1 and −0.4/2 = −0.2. The chain rule carries these on to every weight in θd.

In Python:

import math
m = 2
D_x, D_Gz = [0.9, 0.8], [0.2, 0.4]
# (1/m) Σ [log D(x^(i)) + log(1 − D(G(z^(i))))]
objective = sum(math.log(dx) + math.log(1 - dg) for dx, dg in zip(D_x, D_Gz)) / m
round(objective, 2)  # → -0.53
# its slope with respect to each real note's score: (1 − D) / m
[round((1 - d) / m, 2) for d in D_x]  # → [0.05, 0.1]
# and with respect to each fake's score: −D / m
[round(-d / m, 2) for d in D_Gz]  # → [-0.1, -0.2]

The forger's step: descend this, with respect to its own weights θg.

In words: “average the log of one minus the detective's verdict on each fake, and find which way each of the forger's weights should move to lower that average”, which means raising the verdicts on its fakes.

With the numbers: (log 0.8 + log 0.6) / 2 = −0.367. Its slope with respect to each fake's score is −D / m: −0.1 and −0.2. Descending a negative slope means moving the score up, so the forger changes its fakes until the detective finds them more real. That slope reaches the forger's weights by flowing back through the detective into G.

In Python:

import math
m = 2
D_Gz = [0.2, 0.4]
# (1/m) Σ log(1 − D(G(z^(i))))
round(sum(math.log(1 - d) for d in D_Gz) / m, 3)  # → -0.367
# slope with respect to each fake's score: −D / m
[round(-d / m, 2) for d in D_Gz]  # → [-0.1, -0.2]

Why it matters today

Alternating single steps is still how GANs train, usually with Adam in place of the paper's momentum (see the Adam companion). The paper's warning, that D stays near its best only “so long as G changes slowly enough”, turned out to be the heart of GAN instability. The lesson's train_gan runs this loop on the ring of eight clouds, and shows that the detective's learning rate alone decides whether the forger finds all eight.

The saturation fix · original

“In practice, equation 1 may not provide sufficient gradient for G to learn well.”Goodfellow et al. (2014), §3

Everyday picture

Early on, the forger is terrible and the detective is sure every fake is fake. Picture a game of hotter-and-colder where the helper whispers more quietly the further away you are: exactly when you most need a hint, you hear nothing. The original forger loss behaves like that; it saturates. The paper's fix is to flip the forger's goal from “make the detective less right” to “make the detective say real”, which shouts loudest when the forger is furthest off.

Tiny example

The detective is 99% sure a fake is fake: D(G(z)) = 0.01. The original loss, log(1 − D), changes at a rate of only −0.01 per unit of the detective's score. The flipped goal, maximise log D, written as the loss −log D, changes at −0.99: ninety-nine times more push. When the detective is undecided, D = 0.5, both rates are −0.5.

In words: “the original forger loss moves by only D when the detective's score for a fake moves, so it goes quiet when D is near 0; the flipped loss moves by 1 − D, so it is loudest exactly when the forger is doing worst.” (This pair is our own derivation of the paper's sentence; the paper states the fix in words.)

With the numbers: at D = 0.01 the two rates are −0.01 and −0.99; at D = 0.5 both are −0.5.

In Python:

import math
def sigma(a):
    return 1 / (1 + math.exp(-a))
# a score that makes the detective 99% sure the fake is fake
a = math.log(0.01 / 0.99)
round(sigma(a), 2)  # → 0.01
# measure each slope by nudging the score a tiny bit either way
h = 1e-6
def slope(loss, a):
    return (loss(a + h) - loss(a - h)) / (2 * h)
def original(a):
    return math.log(1 - sigma(a))
def flipped(a):
    return -math.log(sigma(a))
round(slope(original, a), 3)  # → -0.01
round(slope(flipped, a), 3)  # → -0.99
# an undecided detective, score 0 so D = 0.5
round(slope(original, 0.0), 3), round(slope(flipped, 0.0), 3)  # → (-0.5, -0.5)

Hover the chart, or focus it and use the arrow keys, to compare the two pushes at any verdict.

Reading it: the x-axis is the detective's verdict on a fake, on a log scale: far left means “certainly fake”, the right edge “certainly real”. The y-axis is how hard each loss pushes the forger, the size of its slope with respect to the detective's score. The solid line, the original loss, shrinks towards zero on the left, exactly where an untrained forger lives. The dashed line, the flipped loss, stays near 1 there. The two cross at D = 0.5 and share the same finishing point. The paper puts it this way: the same fixed point, much stronger gradients early in learning.

Why it matters today

This footnote-sized paragraph became the default: almost every GAN since uses this non-saturating loss or a Wasserstein-style loss. The lesson's head_start_experiment gives the detective a head start on the ring and measures the forger's gradient under both losses: after 800 detective steps the original loss hands the forger 0.09 while the non-saturating one hands it 15.

4 Theoretical results · original

“The results of this section are done in a non-parametric setting, e.g. we represent a model with infinite capacity by studying convergence in the space of probability density functions.”Goodfellow et al. (2014), §4

Everyday picture

Before trusting a new game, ask two questions. If both players were perfect, where would the game end? And does taking turns actually get there? Section 4 answers both, for idealised players that can be any function, not just a network of a fixed size. It measures the forger by the probability density pg it produces: how thickly its samples cover each spot.

Tiny example

For the whole of §4 we use a world with only three spots, A, B and C. The real data lands on them 60%, 30% and 10% of the time. The forger, still learning, lands on them 10%, 30% and 60% of the time. Everything below can be computed on these six numbers.

Why it matters

These results are why anyone believed GANs could work: the game's only perfect ending is a forger that matches the data. The gap between these idealised players and real networks is where every later GAN paper lives.

4.1 The best detective (Proposition 1) · original

Everyday picture

Suppose that at one spot real notes turn up three times as often as fakes. However clever the detective is, two identical notes look the same to it, so the best it can say there is “75% real”. It should match the local mix, no more and no less.

Tiny example

First the paper rewrites the value as one sum over spots (an integral, since real data is continuous): at each spot, the real share times log D plus the fake share times log(1 − D). On our three spots with a coin-flip detective (D = ½ everywhere) that is −1.386. The paper then notes that at each spot the detective is maximising a · log y + b · log(1 − y), with a the real share and b the fake share, and that this peaks at y = a/(a + b). At a spot with a = 0.3 and b = 0.1, the peak is 0.75: the value is −0.2274 at y = 0.7, −0.2249 at 0.75 and −0.2279 at 0.8.

In words: “walk along every possible x; at each, weigh the log of the verdict by how much real data lives there and the log of one minus the verdict by how many fakes live there; add it all up.” The paper gets here from equation 1 by noticing that averaging over noise z and then pushing it through G is the same as averaging over the fakes' own density pg.

With the numbers: with D = ½ on all three spots, V = (0.6 + 0.3 + 0.1) · log ½ + (0.1 + 0.3 + 0.6) · log ½ = −1.386.

In Python:

import math
# the three-spot world: real and fake shares at A, B, C
p_data = [0.6, 0.3, 0.1]
p_g = [0.1, 0.3, 0.6]
# a coin-flip detective
D = [0.5, 0.5, 0.5]
# ∫ over x becomes a sum over the three spots
V = sum(pd * math.log(d) + pg * math.log(1 - d) for pd, pg, d in zip(p_data, p_g, D))
round(V, 3)  # → -1.386

The step that finishes the proof, for any real share a and fake share b not both zero:

In words: “if a point gets real data with weight a and fakes with weight b, the verdict that earns the detective the most there is the real share, a out of a + b.” The slope of a log y + b log(1 − y) is a/y − b/(1 − y), which is zero exactly at y = a/(a + b).

With the numbers: a = 0.3, b = 0.1: the peak is at 0.3 / 0.4 = 0.75, where the value is −0.2249; its neighbours 0.7 and 0.8 score lower.

In Python:

import math
a, b = 0.3, 0.1
def score(y):
    return a * math.log(y) + b * math.log(1 - y)
# where the peak should be
peak = a / (a + b)
round(peak, 4)  # → 0.75
[round(score(y), 4) for y in (0.7, 0.75, 0.8)]  # → [-0.2274, -0.2249, -0.2279]
# the slope a/y − b/(1 − y) at the peak
round(a / peak - b / (1 - peak), 10)  # → 0.0
0.30 0.10

Reading it: the x-axis is the verdict y the detective could give at one spot, and the curve is what it would earn there, a · log y + b · log(1 − y). It dips steeply at both ends, because a confident mistake is punished by a log that heads to minus infinity, and it has one peak in between. Drag the sliders: the peak always sits at a/(a + b), the readout confirms it, and with a = b it sits at exactly ½. Any detective, however it is built, can do no better at that spot.

So the best detective against a fixed forger is:

In words: “the best verdict at each point is the share of everything found there that is real.”

With the numbers: at A, 0.6 / (0.6 + 0.1) = 0.857; at B, 0.3 / 0.6 = 0.5; at C, 0.1 / 0.7 = 0.143. The detective can be anything at spots where neither real nor fake data ever appears, since nothing there is scored.

In Python:

p_data = [0.6, 0.3, 0.1]
p_g = [0.1, 0.3, 0.6]
# D*(x) = p_data(x) / (p_data(x) + p_g(x)) at A, B, C
D_star = [pd / (pd + pg) for pd, pg in zip(p_data, p_g)]
[round(d, 3) for d in D_star]  # → [0.857, 0.5, 0.143]

Why it matters today

The detective never has to model what real data looks like, only a ratio of two densities. That is why GANs could learn sharp images long before anyone could write down a probability for an image. The lesson's fit_discriminator_table trains the most flexible detective possible and arrives at optimal_discriminator's verdicts without being told the formula.

4.1 Theorem 1: where the game ends · original

“The global minimum of the virtual training criterion C(G) is achieved if and only if pg = pdata. At that point, C(G) achieves the value −log 4.”Goodfellow et al. (2014), §4.1, Theorem 1

Everyday picture

Now let the detective always play its best, and ask what score the forger is left facing. That score, called C(G), turns out to be a fixed number plus a measure of how different the two piles are. The forger's only way to its lowest score is to make the piles identical.

Tiny example

On the three spots, the best detective says 0.857, 0.5 and 0.143, and scores C(G) = −0.990. The coin-flip score −log 4 = −1.386 is the lowest C(G) can ever be. The gap, 0.396, is twice the Jensen-Shannon divergence between the piles, 2 × 0.198: that is how much the forger still has to improve.

In words: “the forger's score is the value with the best detective plugged in: the average log verdict on real data plus the average log of one minus the verdict on fakes, both using D*.” Substituting the formula for D* turns each log into the log of a share, pdata/(pdata + pg) and pg/(pdata + pg).

With the numbers: 0.6 log 0.857 + 0.3 log 0.5 + 0.1 log 0.143 on the real side, and the mirror image 0.1 log 0.143 + 0.3 log 0.5 + 0.6 log 0.857 on the fake side, total −0.990.

In Python:

import math
p_data = [0.6, 0.3, 0.1]
p_g = [0.1, 0.3, 0.6]
D_star = [pd / (pd + pg) for pd, pg in zip(p_data, p_g)]
# E over real x of log D*(x), plus E over fake x of log(1 − D*(x))
C = sum(pd * math.log(d) for pd, d in zip(p_data, D_star)) + sum(pg * math.log(1 - d) for pg, d in zip(p_g, D_star))
round(C, 3)  # → -0.99
# the best C(G) can ever be: log 1/2 + log 1/2
round(-math.log(4), 3)  # → -1.386

To read that gap, the paper uses the Kullback-Leibler (KL) divergence, the extra surprise of expecting q when the truth is p:

In words: “at every x, weigh the log of how much more likely p makes x than q does by how often p produces x, and add up.” It is 0 when p and q agree everywhere and positive otherwise. (The paper uses KL without writing out this definition; we add it so the next line can be checked.)

With the numbers: take p = pdata = (0.6, 0.3, 0.1) and q = the half-and-half mix of the two piles, (0.35, 0.3, 0.35): KL = 0.6 log(0.6/0.35) + 0.3 log 1 + 0.1 log(0.1/0.35) = 0.198.

In Python:

import math
def KL(p, q):
    # Σ_x p(x) log(p(x) / q(x))
    return sum(p_x * math.log(p_x / q_x) for p_x, q_x in zip(p, q) if p_x > 0)
p_data = [0.6, 0.3, 0.1]
p_g = [0.1, 0.3, 0.6]
# the mix (p_data + p_g) / 2
mix = [(pd + pg) / 2 for pd, pg in zip(p_data, p_g)]
mix  # → [0.35, 0.3, 0.35]
round(KL(p_data, mix), 3)  # → 0.198
# a distribution compared with itself
KL(p_data, p_data)  # → 0.0

Subtracting −log 4 from C(G) leaves two KL terms (equation 5), which together are twice the Jensen-Shannon divergence (equation 6):

In words: “the forger's score is the coin-flip floor, −log 4, plus how far the real pile is from the average of the two piles, plus how far the fake pile is from that average; together those two distances are twice the Jensen-Shannon divergence between real and fake.”

With the numbers: −1.386 + 0.198 + 0.198 = −0.990, the same C(G) as before. Jensen-Shannon is never negative and is 0 only when the piles are identical, so C(G) is lowest, −1.386, exactly when pg = pdata. At the other extreme, two piles that never overlap have JSD = log 2 and C(G) = 0.

In Python:

import math
def KL(p, q):
    return sum(p_x * math.log(p_x / q_x) for p_x, q_x in zip(p, q) if p_x > 0)
def JSD(p, q):
    # the average of each pile's KL from the half-and-half mix
    mix = [(a + b) / 2 for a, b in zip(p, q)]
    return (KL(p, mix) + KL(q, mix)) / 2
p_data = [0.6, 0.3, 0.1]
p_g = [0.1, 0.3, 0.6]
round(JSD(p_data, p_g), 3)  # → 0.198
# C(G) = −log 4 + 2 · JSD
round(-math.log(4) + 2 * JSD(p_data, p_g), 3)  # → -0.99
# a forger that matches the data
round(-math.log(4) + 2 * JSD(p_data, p_data), 3)  # → -1.386
# piles with no spot in common
round(JSD([1, 0], [0, 1]), 3), round(math.log(2), 3)  # → (0.693, 0.693)
1.00

Hover the chart, or focus it and use the arrow keys, to read C(G) for any position of the forger.

Reading it: the real pile is a bell curve centred at 0, and the forger's pile is an identical bell curve slid to the position on the x-axis. The solid curve is C(G), what the forger scores against the best detective; the dashed line is the floor, −log 4. At position 0 the piles coincide and C(G) touches the floor, just as Theorem 1 says. Now slide the forger far away: C(G) climbs to 0 and goes flat. Drag the width slider down to make both piles thin, and the flat region spreads almost to the centre. A flat curve has no slope, so a forger out there gets no hint which way to move. The theorem is true, but this flatness is exactly the complaint the Wasserstein GAN paper made three years later (see the Wasserstein GAN companion).

Why it matters today

Theorem 1 is the reason a GAN is said to minimise the Jensen-Shannon divergence. It also gives a reference number: a detective that can do no better than guessing has a loss (−V) of log 4 ≈ 1.386. The lesson measures the flatness directly with js_divergence and compares it with wasserstein_1d.

4.2 Convergence of Algorithm 1 (Proposition 2) · original

“However, the excellent performance of multilayer perceptrons in practice suggests that they are a reasonable model to use despite their lack of theoretical guarantees.”Goodfellow et al. (2014), §4.2

Everyday picture

Picture a bowl. However you drop a marble in, rolling downhill takes it to the single bottom. The proof shows that, if the detective always plays its best, the forger's score C(G) is a bowl in the space of all possible piles (it is convex), with its bottom at pg = pdata. So small downhill steps on the pile itself must end at the data.

Tiny example

The key fact: the largest of several bowls is still a bowl, and its slope at any point is the slope of whichever bowl is on top there. Take two bowls, (x + 1)² and (x − 1)². At x = 2 the first is 9 and the second 1, so the largest is 9, and its slope there is the first bowl's slope, 2 × (2 + 1) = 6. In the proof, each bowl is one detective's score as a function of the pile, and the one on top is the best detective. That is why the forger may simply follow the gradient computed against the best detective.

In words: “define f as the top of a family of convex functions; then at any x, the slope of the function that is on top there is a valid slope for f.”

With the numbers: the family is f−1(x) = (x + 1)² and f+1(x) = (x − 1)². At x = 2, f(2) = max(9, 1) = 9, the top one is β = −1, and its slope 2(2 + 1) = 6 matches the slope of f measured directly.

In Python:

# two convex bowls, one for each α in A = {−1, +1}
def f_alpha(alpha, x):
    return (x - alpha) ** 2
def f(x):
    # f(x) = sup over α of f_α(x)
    return max(f_alpha(alpha, x) for alpha in (-1, 1))
x = 2
f(x)  # → 9
# β = arg sup: which bowl is on top at x
beta = max((-1, 1), key=lambda alpha: f_alpha(alpha, x))
beta  # → -1
# the slope of f_β at x, from the formula 2(x − β)
2 * (x - beta)  # → 6
# the slope of f itself, by nudging x a tiny bit
h = 1e-6
round((f(x + h) - f(x - h)) / (2 * h), 4)  # → 6.0

Hover the chart, or focus it and use the arrow keys, to read the two bowls and the largest of them.

Reading it: the dashed curve is the bowl (x + 1)² and the dotted one is (x − 1)²; the solid line is the largest of the two at every x, drawn over whichever bowl is higher. Left of 0 the second bowl is on top, right of 0 the first. The upper curve is still a bowl, with its bottom at x = 0, and at every point its slope is the slope of whichever bowl it is riding. Replace “bowls” by “detectives” and “x” by “the forger's pile” and you have Proposition 2.

What the proof does not cover

The argument moves the pile pg directly and assumes the detective is at its best at every step. A real GAN moves the weights θg of a network, which is not convex, and the detective only takes k steps. The paper says so plainly, as the quote above shows. Later work found exactly where this breaks: simultaneous gradient steps can circle the equilibrium forever instead of reaching it.

Why it matters today

The gap between Proposition 2 and practice is the story of GAN research. The lesson's dirac_gan shows it with one number per player: plain alternating steps orbit the equilibrium forever, and a gradient penalty on the detective is what makes them spiral in.

5 Experiments · original

“This method of estimating the likelihood has somewhat high variance and does not perform well in high dimensional spaces but it is the best method available to our knowledge.”Goodfellow et al. (2014), §5

Everyday picture

How do you grade a forger that cannot tell you how likely any note is? The paper uses a workaround: drop a small blob of sand on each note the forger makes, then check how much sand lands on real notes it never saw. Lots of sand on the real notes means the forger's output covers the kind of notes that really exist. This is a Parzen window estimate.

Tiny example

The forger has made two samples, at 0 and at 2. Put a bell curve of width σ = 1 on each. A held-out real point at x = 1 sits one unit from both, where each bell has height 0.242, so the average is 0.242 and its log is −1.419. Repeat for every real test point and average the logs: that is the paper's score, a log-likelihood.

In words: “the estimated density at x is the average, over the generator's n samples, of the height at x of a bell curve of width σ centred on each sample.” (The paper describes this procedure in words; the formula is the standard one it refers to.)

With the numbers: samples s = (0, 2), σ = 1, test point x = 1: both bells have height e−1/2 / √(2π) = 0.242 there, so p̂(1) = 0.242 and log p̂(1) = −1.419.

In Python:

import math
def N(x, mean, sigma):
    # the height of a bell curve of width sigma centred on mean
    return math.exp(-0.5 * ((x - mean) / sigma) ** 2) / (sigma * math.sqrt(2 * math.pi))
s, sigma = [0, 2], 1.0
# p̂(x) = (1/n) Σ_i N(x; s_i, σ²) at the test point x = 1
p_hat = sum(N(1, s_i, sigma) for s_i in s) / len(s)
round(p_hat, 3)  # → 0.242
round(math.log(p_hat), 3)  # → -1.419
0.50

Reading it: five illustrative generator samples, at −1.5, −0.3, 0.2, 1.4 and 2.6, each get a bell curve of width σ, and the curve is their average, the Parzen estimate p̂. Two held-out real points sit at 0.8 and 2.0; the readout gives their average log-likelihood. Drag σ down: the estimate becomes a row of spikes, and real points that fall between samples score terribly. Drag it up: the estimate becomes one blur that scores everything alike. The score depends heavily on this one knob, which is why the paper picks σ on a separate validation set, and why it warns that the method is noisy. In 784 dimensions it is far worse, because the gaps between samples are enormous.

What the paper trained and found

  • Datasets: MNIST (handwritten digits), the Toronto Face Database (TFD) and CIFAR-10 (small colour photos).
  • The generator mixed rectified linear and sigmoid units; the detective used maxout units and was trained with dropout (see the dropout companion). Noise entered only at the generator's bottom layer.
Selected rows of Table 1, reproduced with attribution (Goodfellow et al., 2014). Parzen window log-likelihood estimates (higher is better)
ModelMNISTTFD
Stacked CAE121 ± 1.62110 ± 50
Deep GSN214 ± 1.11890 ± 29
Adversarial nets225 ± 22057 ± 26

On MNIST the adversarial nets score highest of the models in the table; on faces a stacked contractive autoencoder (CAE) is slightly ahead. The paper's Figures 2 and 3 show samples rather than numbers: digits, faces and small CIFAR-10 images, each row ending with the nearest real training image to show the model is not copying, and digits made by moving z in a straight line between two noise vectors. The authors claim only that the samples are “at least competitive” with the best generative models of the day.

In Python:

# Table 1: Parzen log-likelihood estimates on MNIST (higher is better)
mnist = {"Stacked CAE": 121, "Deep GSN": 214, "Adversarial nets": 225}
max(mnist, key=mnist.get)  # → 'Adversarial nets'
# how far ahead of the next model, in log-likelihood units
mnist["Adversarial nets"] - mnist["Deep GSN"]  # → 11

Why it matters today

Nobody grades GANs with Parzen windows any more; the measurement problem the paper flagged led to scores that compare whole piles of generated and real images inside a trained image network, most famously FID. The lesson makes the same point on its toy data: a sample-by-sample quality check cannot see a forger that covers only some of the clouds, so mode_coverage counts how many clouds are covered.

6 Advantages and disadvantages · original

“… G must not be trained too much without updating D, in order to avoid ‘the Helvetica scenario’ in which G collapses too many values of z to the same value of x to have enough diversity to model pdata …”Goodfellow et al. (2014), §6

Everyday picture

A forger who finds one note that always fools today's detective is tempted to print nothing else. Each note is convincing, but a pile of identical notes looks nothing like real money. The paper called this the Helvetica scenario; the field now calls it mode collapse.

Tiny example

The real data has four separate clusters. Eight noise values go into the forger. A healthy forger sends two to each cluster: 4 of 4 clusters covered. A collapsed forger sends all eight to the same spot: 1 of 4 clusters covered, and all eight samples look the same, even though each one sits on real data.

Reading it: the lower line holds eight evenly spaced noise values z, and the upper line is the data space, where the four dotted bumps are the real clusters. Each arrow is the generator at work, carrying one z to the sample G(z). Switch to the Helvetica scenario: every arrow converges on one bump. Nothing on the upper line looks wrong in isolation, which is why this failure is easy to miss if you inspect one sample at a time. It happens, as the paper says, when G gets ahead of D: the forger piles onto whatever the current detective likes best, before the detective has learned that this spot is suddenly crowded with fakes.

The trade the paper describes

  • Disadvantages: the model never writes down pg(x), so you cannot ask it how likely an image is; and D must be kept in step with G.
  • Advantages: no Markov chains, only backprop, no inference during learning, and any differentiable function can be the generator.
  • A statistical advantage: the generator is never updated with real examples, only with gradients that come through the detective, so it cannot simply copy pieces of the training data into its weights.
  • Sharpness: a GAN can represent very sharp distributions, even degenerate ones that squeeze all their mass onto a thin sliver of the space, whereas models sampled by a Markov chain need some blur so the chain can move between peaks.

Why it matters today

Both halves of this section held up. Sharp samples were the reason GANs led image generation for years; mode collapse and keeping the players in step were the reasons they were eventually overtaken. The lesson trains the ring of eight clouds three ways and watches the forger hop between alternate clouds when the detective is too slow.

7 Conclusions and future work · original

“A conditional generative model p(x | c) can be obtained by adding c as input to both G and D.”Goodfellow et al. (2014), §7

Everyday picture

Give the forger and the detective the same order slip, “a 20-dollar note” or “the digit 7”, and the forger learns to make whatever the slip asks for, while the detective checks both realism and whether the note matches the slip.

The five directions the paper lists

  1. Conditional models: feed the same extra input c (a label, say) to both G and D, and the forger learns p(x | c).
  2. Learned inference: train a helper network to predict z from x, even after the generator has finished training.
  3. Every conditional at once: model p(xS | the rest of x) for any subset S of the inputs, with networks that share weights.
  4. Semi-supervised learning: reuse the detective's features for a classifier when few labels are available.
  5. Efficiency: better ways to coordinate G and D, or better distributions to draw z from.

Why it matters today

The last item, coordinating G and D, turned out to be the hard one: separate learning rates, gradient penalties, spectral normalization and the Wasserstein loss all attack it. The GAN lesson runs each of these fixes and shows what it does.

What changed since 2014

The two-player game is unchanged. Nearly everything around it has been replaced:

Choice in the paperCommon laterWhy
Minimax loss, with the non-saturating version as a practical tipNon-saturating or Wasserstein-style losses by defaultA forger that is far off still gets a strong gradient
Fully connected networks with maxout (one convolutional CIFAR-10 model)Convolutional generators and detectives (DCGAN), later style-based generators (StyleGAN)Images have local structure that convolutions share across positions
SGD with momentum, k = 1Adam with low momentum (β1 = 0.5), separate learning rates per playerThe target keeps moving, so less momentum and a faster detective help
Jensen-Shannon, through the log-loss gameWasserstein distance (companion), gradient penalties, spectral normalizationJS goes flat when real and fake don't overlap
Parzen window log-likelihoodScores that compare whole piles of images, such as FIDParzen estimates are unreliable in high dimensions
GANs as the image generatorDiffusion models generate; adversarial losses sharpen decoders and distil fast samplersDiffusion trains stably and covers every mode

The diffusion lesson covers what took over, and the closing section of the GAN lesson shows where adversarial losses live on.

Glossary

Every term with hover guidance on this page, in one place.