An annotated companion · AI Primer

Dropout: A Simple Way to Prevent Neural Networks from Overfitting, annotated

Paper at JMLR PDF
About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. Equations are reproduced with every symbol decoded. The paper's figures are redrawn from scratch, and only a few selected numbers from its tables are shown, each with attribution. Read the original alongside: every section heading links to the page of the PDF where that section begins.

How to read this page

  • Any dotted word explains itself on hover, focus or tap; so does every symbol in every equation.
  • The network in §1 is live: press Resample to watch a new “thinned” network appear at every training step.

Each idea climbs the ladder: everyday picture, tiny example, diagram, math, why it matters today. The regularization lesson implements dropout from scratch and measures what it does to overfitting.

Abstract · original

“The key idea is to randomly drop units (along with their connections) from the neural network during training.”Srivastava et al. (2014), Abstract

Everyday picture

A sports team that trains every day with a few random players benched learns not to depend on any single star: every player must be useful with any combination of teammates. On match day everyone plays. Dropout does this to a neural network: during training each neuron is switched off at random for each example, and at test time all of them work together, slightly turned down to compensate.

What the paper claims

  • Large networks overfit, and averaging many networks would help but is too expensive. Dropout gets most of that benefit from one network.
  • Training with dropout is like training an exponential number of “thinned” networks that share weights; using the full network with scaled-down weights approximates averaging them all.
  • It improved results in vision, speech, text and computational biology, reaching the best published results on several benchmarks.

Why it matters today

Dropout became a default ingredient for a decade, and the original Transformer used it throughout. Very large language models trained on enormous datasets often use little or none, but the idea behind it, that injecting noise during training makes models robust, runs through modern deep learning.

1 Introduction · original

Everyday picture

With limited training data, a big network learns real patterns and accidents of the sample alike. The gold standard fix is to ask many different models and average their answers, but training and running many big networks is prohibitively expensive. Dropout's trick is to train a single network in a way that makes it behave like a crowd.

Tiny example

A network with n droppable units can be thinned in 2n ways, since each unit is either kept or dropped. With 5 units that is 25 = 32 sub-networks; with the paper's large MNIST network of 2 layers × 8,192 units it is 216,384, a number with about 4,900 digits. They all share the same weights, so storage stays that of one network. Each training example trains one freshly sampled thinned network, so most of them are trained rarely or never, yet sharing weights lets them all benefit.

Reading it: the network has 4 inputs (pink), two hidden layers of 5 units (blue) and 2 outputs (green). Each press of Resample is one training step: every hidden unit is kept with probability p (here 0.5, the paper's default for hidden units), and dropped units, crossed out in red, vanish with all their connections. What remains is the thinned network trained on this one example. Press it a few times and notice that no two steps train the same network. Lower p and more units drop. Tick test time and every unit is back, with the readout showing the outgoing weights multiplied by p, as §4 explains. (The paper usually keeps inputs with p = 0.8; this demo drops only hidden units, to keep the picture simple.)

Why it matters today

“Many models for the price of one” is the recurring theme: shared weights, sampled sub-networks, one cheap approximation at test time. The same idea reappears in stochastic depth (dropping whole layers) and in using dropout at test time to estimate uncertainty.

2 Motivation · original

“Ten conspiracies each involving five people is probably a better way to create havoc than one big conspiracy that requires fifty people to all play their parts correctly.”Srivastava et al. (2014), §2

Everyday picture

The paper offers two analogies. Evolution: sexual reproduction splits up sets of genes that work well together, which should hurt, yet most advanced life reproduces that way. One theory is that it rewards genes that are useful in any company, which makes the whole organism more robust. Conspiracies: a plan that needs fifty people to act perfectly is fragile, while many small plans survive surprises. In a network, a neuron that only works when specific other neurons fix its mistakes is part of a fragile fifty-person plan. Dropout makes such co-adaptation impossible, because any partner might be missing.

Tiny example

If a neuron's usefulness depends on 3 specific partners, each present with probability 0.5, all three are present together only 0.53 = 12.5% of the time. A neuron that relies on that team is useless 87.5% of the time, so training pushes it towards features that help on their own.

Why it matters today

“Break up co-adaptation” remains the most common one-line intuition for why dropout works, alongside “it is a cheap ensemble”. The paper checks the first intuition directly in §7.1.

3 Related work · original

Everyday picture

Adding noise during training was not new: denoising autoencoders corrupt their inputs and learn to reconstruct the clean version. Dropout applies noise to hidden units as well, frames it as model averaging, and uses it for ordinary supervised learning.

Tiny example

Denoising autoencoders typically worked best with about 5% input noise. Thanks to the test-time weight scaling of §4, dropout tolerated far more: dropping 20% of input units and 50% of hidden units was often best.

Why it matters today

“Corrupt the input and learn to repair it” later became the training objective of masked language models such as BERT: a close cousin of the denoising line this section cites.

4 Model description · original

Everyday picture

Before each layer reads its inputs, flip a biased coin for every unit in the layer below: heads (probability p) the unit speaks, tails it stays silent. The layer then does its usual weighted sum over whatever it heard.

Tiny example

A layer below outputs y = (2, 0.5, 1). The coins come up r = (1, 0, 1), so the thinned outputs are ỹ = (2, 0, 1). With incoming weights w = (0.5, 0.4, −1) and bias 0.1, the summed input is 0.5 × 2 + 0.4 × 0 − 1 × 1 + 0.1 = 0.1, and ReLU gives 0.1. With a different mask, r = (1, 1, 0), it would be 1 + 0.2 + 0 + 0.1 = 1.3.

In words: “for every unit in layer l, draw a 1 with probability p and a 0 otherwise; multiply the layer's outputs by those 0s and 1s; the next layer computes its usual weighted sum and nonlinearity on what survives. At test time, keep every unit but multiply the weights by p.”

With the numbers: y = (2, 0.5, 1) and r = (1, 0, 1) give ỹ = (2, 0, 1); z = 0.5·2 + 0.4·0 − 1·1 + 0.1 = 0.1. At test time with p = 0.5 the weights become (0.25, 0.2, −0.5) and z = 0.5 + 0.1 − 0.5 + 0.1 = 0.2.

In Python:

# y^(l), the layer's outputs
y = [2, 0.5, 1]
# r^(l), one Bernoulli(p) draw per unit
r = [1, 0, 1]
# ỹ = r * y
y_tilde = [r_j * y_j for r_j, y_j in zip(r, y)]
y_tilde  # → [2, 0.0, 1]
# w_i^(l+1), b_i^(l+1)
w, b = [0.5, 0.4, -1], 0.1
z = sum(w_j * y_j for w_j, y_j in zip(w, y_tilde)) + b
round(z, 2)  # → 0.1
p = 0.5
# W_test = p W
W_test = [p * w_j for w_j in w]
W_test  # → [0.25, 0.2, -0.5]
# every unit kept
round(sum(w_j * y_j for w_j, y_j in zip(W_test, y)) + b, 2)  # → 0.2

Why scale the weights at test time?

Everyday picture: a choir rehearses with half its singers absent at random each time, so the audience has come to expect a half-sized choir. On concert night everyone sings; if they all sing at full voice it is twice as loud as rehearsal. So everyone sings at half volume. Multiplying the weights by p does exactly that.

training test present withprobability p w alwayspresent p·w next layer receives, on average, p·w·y

Hover or tap a part.

Figure 2 of the paper, redrawn: a unit at training time and at test time. Based on Srivastava et al. (2014), Figure 2.

Reading it: on the left, during training, the unit takes part only a fraction p of the time; when present it sends its output y through weight w. On average the next layer receives p × w × y. On the right, at test time, the unit is always present, so its weight is scaled to p × w: the next layer receives exactly p × w × y, what it was used to receiving on average. With y = 2, w = 0.5 and p = 0.5: during training, half the time 1.0 and half the time 0, averaging 0.5; at test time 0.25 × 2 = 0.5.

Why it matters today

Frameworks now use inverted dropout, which the paper notes in §10 is equivalent: divide the surviving outputs by p during training, and use the weights unchanged at test time. Test-time code then does not need to know dropout exists.

5 Learning dropout nets · original

Everyday picture

Training is ordinary backpropagation, run on each example's thinned network. A weight that belonged to a dropped unit simply gets zero gradient from that example. The paper pairs dropout with one extra rule that it found especially helpful: a max-norm constraint, a leash on how long each neuron's weight vector can get.

Tiny example

With c = 3, a neuron's weights w = (3, 4) have length 5, too long, so they are rescaled to 3/5 of their size: (1.8, 2.4), which has length exactly 3. The leash lets training use a large learning rate without weights blowing up, while dropout's noise keeps it exploring.

In words: “the length of every hidden unit's incoming weight vector must stay at or below c.”

With the numbers: ‖(3, 4)‖ = √(9 + 16) = 5 > 3, so project: (3, 4) × 3/5 = (1.8, 2.4).

In Python:

import math
w, c = [3, 4], 3
# ‖w‖₂
norm = math.sqrt(sum(w_j ** 2 for w_j in w))
norm, norm <= c  # → (5.0, False)
# project back onto ‖w‖₂ = c
[round(w_j * c / norm, 2) for w_j in w]  # → [1.8, 2.4]

The paper's recipe: dropout plus max-norm, large decaying learning rates and high momentum. For networks pretrained with RBMs or autoencoders (§5.2), the pretrained weights are first scaled up by 1/p and fine-tuned with smaller learning rates, so the noise does not wipe out what pretraining learned.

Why it matters today

Max-norm faded as normalization layers and weight decay took over, but the underlying worry, keeping weights bounded so aggressive learning rates stay safe, is exactly what those replacements address.

6 Experimental results · original

Everyday picture

If dropout is a general tool and not a trick for one dataset, it should help everywhere. The authors test images, speech, text and genetics. Every dataset improved.

Selected results from §6 of Srivastava et al. (2014): error rates (%) without and with dropout, lower is better
Dataset (§)What it isWithoutWith dropoutNote
MNIST (6.1.1)handwritten digits1.600.95best plain net → dropout, ReLU, max-norm, 2 × 8,192 units
SVHN (6.1.2)house numbers from Street View3.952.55dropout in all layers, convolutional ones included
CIFAR-10 (6.1.3)tiny photos, 10 classes14.9812.61no data augmentation
CIFAR-100 (6.1.3)tiny photos, 100 classes43.4837.20“a huge improvement”
ImageNet 2012 (6.1.4)1,000-class photos≈ 2616.4top-5 test error; the best non-neural entries versus conv nets with dropout, which won ILSVRC 2012
TIMIT (6.2)speech, phone error rate23.421.86-layer net
Reuters (6.3)news topics from word counts31.0529.62a much smaller gain for text

Reading it: each bar is the relative reduction in error from adding dropout, computed from the table above: (without − with) / without. Vision gains are large, from about 16% on CIFAR-10 to about 41% on MNIST, where dropout also allowed a much bigger network; speech gains are moderate; text is smallest at about 5%. The paper notes this pattern too: the bag-of-words text model benefited least.

Two careful comparisons

  • Against Bayesian neural networks (6.4): on a small genetics dataset, a properly Bayesian network, which weighs every model by how well it fits, still scored best (code quality 623 bits, higher is better), but dropout (567) beat everything else, including standard networks (440). Dropout is an equal-weights shortcut to the Bayesian ideal, far cheaper to train and use.
  • Against other regularizers (6.5): on the same MNIST network, L2 weight decay gave 1.62% error, max-norm alone 1.35%, dropout with L2 1.25%, and dropout with max-norm 1.05%.

Why it matters today

The ImageNet row is AlexNet (Krizhevsky et al., 2012): the network that started the deep-learning boom used dropout in its fully connected layers. The companion to that story is the CNN and RNN lesson.

7 Salient features · original

Everyday picture

Five experiments open the box. Each asks one question about what dropout actually does inside a network.

  1. Features (7.1). An autoencoder with 256 hidden units trained without dropout learns features that look like noise individually: they only work as a tangled team. With dropout (p = 0.5), individual units learn recognisable edges, strokes and spots. This is co-adaptation, broken, made visible.
  2. Sparsity (7.2). Dropout makes activations sparse without being asked to: the average hidden activation fell from about 2.0 to about 0.7, with far more units near zero.
  3. Dropout rate (7.3). With the architecture fixed, very small p underfits; test error is flat for p between 0.4 and 0.8 and rises again as p approaches 1. The default p = 0.5 is close to best.
  4. Data size (7.4). On 100 or 500 examples dropout did not help: the network could memorise even through the noise. The gain grows with more data, up to a point, then shrinks, because with enough data overfitting stops being the problem.
  5. Averaging (7.5). Sampling k thinned networks and averaging their predictions matches the cheap weight-scaling rule at around k = 50, and is only slightly better beyond. So weight scaling is a good approximation.

Try it: Monte-Carlo averaging versus weight scaling

2 hidden layers of 64 ReLU units, p = 0.5, one fixed input

Hover or tap to compare the estimates at each k.

Reading it: this is a live toy, not the paper's experiment: a small random network in your browser, one input, and the probability it assigns to class 1. Every line is plotted as its difference from the true average over thinned networks (estimated with 20,000 random masks), so the flat green dotted line at 0 is the truth. The flat orange dashed line is the weight-scaling answer: one pass with every unit on and weights multiplied by p. The wobbly solid blue line is the Monte-Carlo estimate from the first k masks. At small k it jumps around; by one or two hundred masks it has settled close to the green line (the paper's trained MNIST network needed about 50 to match weight scaling). The orange and green lines are close but not identical: weight scaling is exact for a single linear layer and an approximation once nonlinearities are stacked, which matches the paper's finding that it is “a fairly good approximation”.

Why it matters today

Keeping dropout on at test time and averaging several passes, which the paper calls Monte-Carlo model averaging, was later reinterpreted as an approximate Bayesian method for estimating a model's uncertainty (“MC dropout”, Gal and Ghahramani, 2016).

8 Dropout restricted Boltzmann machines · original

Everyday picture

A restricted Boltzmann machine is a two-layer model that learns what typical inputs look like. The dropout trick carries over: randomly silence hidden units while it learns. Dropout RBMs learned coarser features, had very few dead units, and produced much sparser activity than standard RBMs.

Why it matters today

RBMs themselves have faded, but this section makes the paper's broader point: dropout is a general principle (train a large model by sampling sub-models from it), not a trick for one architecture.

9 Marginalizing dropout · original

Everyday picture

Instead of flipping coins, you could work out the average effect of all the coin flips with algebra and train on that directly: no randomness, same result on average. For linear regression this is possible, and the answer is a familiar friend: dropout turns into a form of L2 regularization (ridge regression).

Tiny example (check it by hand)

One data point x = (1, 2) with target y = 3, weights w = (0.5, 0.5), and p = 0.5. The four equally likely masks give predictions 0, 0.5, 1 and 1.5, so squared errors of 9, 6.25, 4 and 2.25, averaging 5.375. The formula below gives (3 − 0.5 × 1.5)2 + 0.5 × 0.5 × (1² × 0.5² + 2² × 0.5²) = 5.0625 + 0.3125 = 5.375. They agree.

In words: “the average squared error over all dropout masks equals the squared error of a model whose inputs are scaled by p, plus a penalty on the weights; each weight's penalty is scaled by how much its input varies.”

With the numbers: ‖3 − 0.5 × (0.5 + 1)‖² = 2.25² = 5.0625; Γ = diag(1, 2) for this single point, so ‖Γw‖² = 0.25 + 1 = 1.25 and p(1 − p) × 1.25 = 0.3125.

In Python:

from itertools import product
x, y, w, p = [1, 2], 3, [0.5, 0.5], 0.5
# every R; at p = 0.5 all four are equally likely
masks = list(product([0, 1], repeat=2))
sum((y - sum(R_j * x_j * w_j for R_j, x_j, w_j in zip(R, x, w))) ** 2 for R in masks) / len(masks)  # → 5.375
# ‖y − pXw‖²
fit = (y - p * sum(x_j * w_j for x_j, w_j in zip(x, w))) ** 2
# p(1 − p)‖Γw‖², Γ = diag(1, 2)
penalty = p * (1 - p) * sum((x_j * w_j) ** 2 for x_j, w_j in zip(x, w))
fit, penalty, fit + penalty  # → (5.0625, 0.3125, 5.375)

The penalty's strength is p(1 − p): largest at p = 0.5 and vanishing as p → 1 (no dropout). And because Γ holds each input's spread, dropout squeezes the weights on high-variance inputs hardest. For logistic regression and deep networks there is no exact closed form, though approximate versions exist.

Why it matters today

This is the cleanest proof that dropout is a regularizer and not magic: in the simplest case it is weight decay, with a data-dependent strength. See the regularization lesson for L1, L2 and dropout side by side.

10 Multiplicative Gaussian noise · original

Everyday picture

Dropout multiplies each activation by a random 0 or 1. Why not multiply by a random number from a smooth bell curve centred on 1? The paper found that this “Gaussian dropout” works as well or slightly better, and it needs no rescaling at test time, because the average multiplier is already 1.

Tiny example

Inverted Bernoulli dropout with p = 0.5 multiplies by 2 half the time and 0 half the time: mean 1, variance (1 − p)/p = 1. Gaussian dropout with σ2 = (1 − p)/p = 1 multiplies by a draw from N(1, 1): the same mean and variance, but the multiplier could be 0.3, 1.7 or −0.4.

In words: “multiply each activation by a random number drawn from a bell curve with mean 1, choosing its variance to match the Bernoulli dropout it replaces.”

With the numbers: p = 0.8 gives σ2 = 0.2 / 0.8 = 0.25, so σ = 0.5. In the paper's Table 10, Gaussian dropout gave 0.95% error on MNIST against 1.08% for Bernoulli, and 12.5% against 12.6% on CIFAR-10 (averaged over 10 random seeds).

In Python:

import math
p = 0.8
# σ² = (1 − p) / p
sigma2 = (1 - p) / p
round(sigma2, 2), round(math.sqrt(sigma2), 2)  # → (0.25, 0.5)

Why it matters today

The observation that the shape of the noise matters less than its mean and variance opened the door to many variants, and to interpreting dropout as approximate Bayesian inference.

11 Conclusion · original

“A dropout network typically takes 2-3 times longer to train than a standard neural network of the same architecture.”Srivastava et al. (2014), §11

The honest cost: every step trains a different random architecture, so the gradients are noisy and training takes two to three times longer. That noise is also why it works. The trade-off is overfitting against training time, and the paper leaves “speeding up dropout” as future work.

What changed since 2014

In the paperCommon todayWhyWhere to learn more
Scale weights by p at test timeInverted dropout: divide by p during trainingTest-time code needs no dropout logicregularization lesson
p = 0.5 in hidden layers everywhereRates of 0.1 or less (keep 0.9+), often only in specific placesBatch/layer normalization and larger datasets already regularizeTransformer companion (dropout 0.1)
Dropout plus max-normDropout plus weight decay (AdamW)Normalization layers and decoupled weight decay took over max-norm's jobAdam companion
Monte-Carlo averaging as a checkMC dropout for uncertainty estimatesThe spread of predictions across masks signals how unsure the model is
Essential for large networksOften zero for large language model pretrainingWith trillions of tokens seen roughly once, there is little to overfit toGPT-3 companion

Glossary

Every term with hover guidance on this page, in one place.