Adam: A Method for Stochastic Optimization, annotated
How to read this page
Nothing here assumes you already know the jargon.
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and every equation is followed by a table of its symbols.
- The diagrams and demos are live. The optimizer race in §2.1 is the fastest way to feel what Adam does.
Each idea climbs the same ladder: an everyday picture, a tiny example you can check by hand, a diagram, the math, and why it matters today. The optimizers lesson builds SGD, momentum, Adam and AdamW from scratch in NumPy.
Abstract · original
“We introduce Adam, an algorithm for first-order gradient-based optimization of stochastic objective functions, based on adaptive estimates of lower-order moments.”Kingma and Ba (2014), Abstract
Everyday picture
Training a neural network means walking downhill on a landscape made of the model's error (the loss), in thick fog, feeling the slope under your feet (the gradient). The question every optimizer answers is: how big a step should I take, in which direction? Adam's answer: keep a running average of the slopes you have felt (so a single bumpy step doesn't throw you), keep a running average of how large the slopes have been, and take each step in proportion to their ratio. Where the ground has been steep and erratic you tread carefully; where it has been gentle but consistent you stride out. And it does this separately for every one of the model's millions of adjustable numbers.
What the paper claims
- A method that needs only gradients (“first-order”), little memory, and is simple to implement.
- Step sizes that do not change if every gradient is multiplied by a constant, and that are roughly capped by one setting, the step size α.
- Hyperparameters that “typically require little tuning”, with defaults that became universal: α = 0.001, β1 = 0.9, β2 = 0.999, ε = 10−8.
- A convergence guarantee for convex problems, and experiments where Adam matches or beats the alternatives on logistic regression, fully connected nets and convolutional nets.
Why it matters today
Adam, and its descendant AdamW, is the default optimizer for Transformers, including every large language model. If you train anything with a neural network today, this paper is probably running inside your training loop.
1 Introduction · original
“The method computes individual adaptive learning rates for different parameters from estimates of first and second moments of the gradients; the name Adam is derived from adaptive moment estimation.”Kingma and Ba (2014), §1
Everyday picture
Imagine tuning a mixing desk with a thousand sliders, where some sliders make a huge difference with a tiny nudge and others barely matter. A single “how far to move” setting for all of them is hopeless: too big and the sensitive sliders overshoot, too small and the insensitive ones never get anywhere. Adam gives every slider its own step size, learned from that slider's own history. The paper names the two things it tracks after statistics: the first moment (the average) and the second moment (the average of the squares, a measure of size).
Tiny example
Suppose weight A's recent gradients were (10, 12, 8) and weight B's were (0.1, 0.12, 0.08). Plain stochastic gradient descent with learning rate 0.01 moves A by about 0.1 and B by about 0.001: a hundredfold difference, purely because of scale. Adam divides each average gradient by the typical size of that weight's gradients. For A that is roughly 10 / 10 = 1, for B roughly 0.1 / 0.1 = 1, so both move by about the same step, α.
Two ancestors it combines
- AdaGrad works well when gradients are sparse.
- RMSProp works well when the problem keeps changing (“non-stationary”).
Adam takes RMSProp's running average of squared gradients, adds a running average of the gradients themselves (momentum), and fixes a start-up flaw both share (§3).
Why it matters today
Per-weight step sizes are exactly what deep networks need: the gradients of an embedding table, an attention matrix and a layer-norm gain differ by orders of magnitude. The optimizers lesson races these methods on an ill-conditioned bowl.
2 The algorithm · original
Everyday picture
Think of a hiker with two notebooks. In the first, after every step, they update a running average of which way the ground slopes: mostly the old average, a little of the newest reading. In the second, they update a running average of how steep the ground has been, ignoring direction. Before each step they divide the first by the square root of the second: “which way, relative to how steep it usually is.” That ratio, times a fixed stride length α, is the step.
Tiny example: three steps on one weight
One weight θ starts at 1.0. Settings: α = 0.1, β1 = 0.9, β2 = 0.999, ε = 10−8. The gradients arrive as 2, then 1, then 1.5.
| t | gt | mt | vt | m̂t | v̂t | step α·m̂/√v̂ | θt |
|---|---|---|---|---|---|---|---|
| 1 | 2 | 0.1 × 2 = 0.2 | 0.001 × 4 = 0.004 | 0.2 / 0.1 = 2 | 0.004 / 0.001 = 4 | 0.1 × 2 / 2 = 0.1 | 0.9 |
| 2 | 1 | 0.9 × 0.2 + 0.1 × 1 = 0.28 | 0.999 × 0.004 + 0.001 × 1 = 0.004996 | 0.28 / 0.19 = 1.474 | 0.004996 / 0.001999 = 2.499 | 0.1 × 1.474 / 1.581 = 0.0932 | 0.8068 |
| 3 | 1.5 | 0.402 | 0.007241 | 1.483 | 2.416 | 0.0954 | 0.7114 |
Notice that the first step is exactly α = 0.1, and later steps stay close to it even though the gradients vary. That is the “step size roughly capped at α” property of §2.1, visible in three rows.
Hover or tap a box. Start with the gradient at the bottom and follow the arrows up.
Reading it: read from the bottom up; this loop runs once per training step, separately for every weight. The fresh gradient splits into two lanes. The left lane (blue) averages the gradient itself into m, which gives the direction. The right lane (orange) averages the squared gradient into v, which gives the typical size of the gradient. Both averages start at zero, so the yellow boxes rescale them to undo that start-up bias (§3). The right lane is square-rooted to get back to gradient units, the left lane is divided by it, and the result times α is subtracted from the weight. The dashed box marks the optimizer's memory: m and v are two extra numbers stored for every weight in the model.
The math
In words: “get the gradient; blend it into a running average of gradients and its square into a running average of squared gradients; scale both averages up to undo their zero start; then move the weight by α times the average gradient divided by the square root of the average squared gradient.”
With the numbers: at t = 2 in the table: m2 = 0.9 × 0.2 + 0.1 × 1 = 0.28, v2 = 0.999 × 0.004 + 0.001 × 1² = 0.004996, m̂2 = 0.28 / (1 − 0.81) = 1.474, v̂2 = 0.004996 / (1 − 0.998001) = 2.499, so θ2 = 0.9 − 0.1 × 1.474 / 1.581 = 0.8068.
In Python:
import math
alpha, beta1, beta2, eps = 0.1, 0.9, 0.999, 1e-8
# row t = 1 of the table
m_prev, v_prev, theta_prev = 0.2, 0.004, 0.9
# step 2, gradient g_2 = 1
t, g = 2, 1
# m_t = β1 m_(t−1) + (1 − β1) g_t
m = beta1 * m_prev + (1 - beta1) * g
# v_t = β2 v_(t−1) + (1 − β2) g_t²
v = beta2 * v_prev + (1 - beta2) * g ** 2
round(m, 2), round(v, 6) # → (0.28, 0.004996)
# m̂_t = m_t / (1 − β1^t)
m_hat = m / (1 - beta1 ** t)
# v̂_t = v_t / (1 − β2^t)
v_hat = v / (1 - beta2 ** t)
round(m_hat, 3), round(v_hat, 3) # → (1.474, 2.499)
# θ_t = θ_(t−1) − α m̂_t / (√v̂_t + ε)
theta = theta_prev - alpha * m_hat / (math.sqrt(v_hat) + eps)
round(theta, 4) # → 0.8068
Why it matters today
Every operation is element-wise: no matrix inverses, no second derivatives. That is why Adam scales to billions of weights. The price is memory: two extra numbers (m and v) per weight. For a 7-billion-parameter model with 32-bit optimizer state, that is 2 × 7 × 109 × 4 bytes = 56 GB before you store the weights themselves.
2.1 Adam's update rule · original
“This can be understood as establishing a trust region around the current parameter value, beyond which the current gradient estimate does not provide sufficient information.”Kingma and Ba (2014), §2.1
Everyday picture
A careful driver in fog never goes faster than they can see. Adam's steps are capped the same way: whatever the gradients look like, a step is usually no bigger than α. And the ratio m̂ / √v̂ behaves like a signal-to-noise ratio: when the gradients agree with each other it is close to ±1 and Adam strides; when they point every which way, the average cancels out, the ratio shrinks, and Adam tiptoes. Near the bottom of the valley the gradients become noisy, so the steps shrink automatically.
Tiny example: scale invariance
Multiply every gradient in the §2 table by 100 (as if the loss were measured in cents instead of dollars). Then m̂ grows 100×, v̂ grows 100² = 10,000×, and √v̂ grows 100×, so the ratio is unchanged: the first step is still exactly 0.1 and the second still 0.0932. Plain SGD's step would grow 100×.
In words: “the step is α times the ratio of the average gradient to its typical size; that ratio is usually between −1 and 1, so the step is usually at most α; and multiplying every gradient by c changes the top and bottom of the ratio equally, so it cancels.”
With the numbers: with c = 100 at t = 1: 100 × 2 / √(100² × 4) = 200 / 200 = 1, the same ratio as 2 / √4 = 1, so the step stays 0.1 × 1 = 0.1.
In Python:
import math
alpha, c = 0.1, 100
# t = 1 in the §2 table
m_hat, v_hat = 2, 4
# (c m̂_t) / √(c² v̂_t): every gradient times c
c * m_hat / math.sqrt(c ** 2 * v_hat) # → 1.0
# m̂_t / √v̂_t: the same ratio, c cancelled
m_hat / math.sqrt(v_hat) # → 1.0
# Δ_t = α · m̂_t / √v̂_t
alpha * m_hat / math.sqrt(v_hat) # → 0.1
Try it: the optimizer race
Click anywhere on the surface to choose where all three start, then press Run.
Reading it: the picture is a map of the loss seen from above; the rings are contour lines and the minimum is the centre. The valley is 25 times steeper across (up and down) than along (left to right), which is the typical shape of a real network's loss. Press Run. SGD (blue) has to use a small step, because a bigger one would overshoot the steep walls: it drops down the wall at once, then crawls along the nearly flat floor. Momentum (orange) builds speed but swings from wall to wall. Adam (green) rescales each direction separately, so it takes similar-sized steps across and along the valley, cutting diagonally towards the centre before settling, and ends with the lowest loss. Now drag loss scale c, which multiplies the whole loss (and so every gradient) by c. Down at 0.1, SGD and momentum barely move. Above about 3, SGD overshoots the steep walls and flies off; above about 30, momentum does too. Adam's path does not change at all: watch its end point in the readout stay at the same coordinates for every c. That is the scale invariance from the equation above. Tick noisy gradients to see Adam's steps shrink near the minimum as the signal-to-noise ratio falls. This is a toy two-weight problem computed live in your browser, not a result from the paper.
Why it matters today
Because the step is roughly bounded by α, you can often guess a sensible α before training starts. The shrinking-near-the-optimum behaviour is also why Adam is usually paired with a learning-rate schedule: the schedule handles the global pace, Adam handles the per-weight balance.
3 Initialization bias correction · original
“In algorithm 1 we therefore divide by this term to correct the initialization bias.”Kingma and Ba (2014), §3
Everyday picture
You start a running average of daily temperatures with an empty notebook, so you write down “0°” as yesterday's average. After one warm day of 20°, a 90%-old / 10%-new average says 2°, which is absurd. Your average is dragged towards the made-up zero until enough real days have washed it out. The fix is to divide by the fraction of the average that is real data so far: after one day that fraction is 10%, and 2° / 0.1 = 20°, which is right.
Tiny example
In the §2 table, v1 = 0.004 although the squared gradient was 4. Only 1 − 0.9991 = 0.1% of v is real data, so v̂1 = 0.004 / 0.001 = 4. Without the correction the first step would be 0.1 × 0.2 / √0.004 = 0.316, more than 3× too big, and by step 12 the uncorrected step would be about 6.6× too big. For v, with β2 = 0.999, it takes about 1,000 steps before the correction stops mattering.
In words: “unrolled, v is a weighted sum of all past squared gradients whose weights add up to only 1 − β2t, not 1; so on average v is the true squared gradient shrunk by that factor, and dividing by it undoes the shrinkage.”
With the numbers: after t = 10 steps, the weights in v add up to 1 − 0.99910 = 0.00996, so v is about 1% of the truth; after 1,000 steps they add up to 1 − 0.9991000 = 0.632.
In Python:
beta2, t = 0.999, 10
# the weight on each g_i² inside Σ
weights = [(1 - beta2) * beta2 ** (t - i) for i in range(1, t + 1)]
# they add up to 1 − β2^t, not 1
round(sum(weights), 5) # → 0.00996
round(1 - beta2 ** t, 5) # → 0.00996
# after 1,000 steps
round(1 - beta2 ** 1000, 3) # → 0.632
Hover or tap the chart to read values at any step.
Reading it: the x-axis is the training step on a log scale. The two lower curves show how much of m and of v is real data so far (1 − β1t and 1 − β2t); both start near zero and climb to 1. The top curve is how much bigger each step would be without bias correction: the ratio (1 − β1t) / √(1 − β2t). With the default β2 = 0.999 it peaks at about 6.6× at step 12. Slide β2 to 0.9999, as sparse problems need, and the peak grows to about 21×. The paper's §6.4 found exactly this: without the correction, β2 close to 1 made training unstable early on.
Why it matters today
This is the one ingredient that separates Adam from “RMSProp with momentum”, and it is what makes the first few hundred steps of training safe. Learning-rate warmup, now standard for Transformers, attacks the same early-training fragility from a different side.
4 Convergence analysis · original
Everyday picture
Picture a weather forecaster who has to commit to a forecast each morning before seeing the day's weather, and is scored against the single best fixed forecast they could have chosen with hindsight. The shortfall, added up over all days, is called regret. A good strategy's regret grows more slowly than the number of days, so its average shortfall per day shrinks to zero: you end up as good as the best fixed choice.
Tiny example
If regret grows like √T, then after 100 days the average shortfall is proportional to √100 / 100 = 0.1, and after 10,000 days to √10000 / 10000 = 0.01. More days means a smaller average gap, heading to zero.
In words: “regret adds up, over every step, how much worse Adam's current weights did on that step's data than the best fixed weights would have; the paper proves this total grows no faster than the square root of the number of steps, so the per-step average vanishes.”
With the numbers: √T / T = 1/√T, so at T = 1,000,000 steps the average regret is on the order of 1/1000 of its scale.
In Python:
import math
T = 1_000_000
# R(T) / T when R(T) grows like √T
math.sqrt(T) / T # → 0.001
# the same thing: 1 / √T
1 / math.sqrt(T) # → 0.001
The proof needs assumptions that real networks break: the loss must be convex, gradients bounded, the step size decaying like α/√t, and β1 decaying over time. The paper is explicit that the analysis does not cover non-convex problems.
Why it matters today
In 2018 Reddi, Kale and Kumar showed a flaw in this proof and simple convex problems on which Adam does not converge (their fix is called AMSGrad). In practice this rarely bites, and nobody picks Adam for its theorem: it won on its behaviour in experiments, the subject of §6.
5 Related work · original
Everyday picture
Three hikers, three note-taking habits. The AdaGrad hiker writes down every squared slope forever and divides by the square root of the running total, so they get more and more cautious and eventually barely move. The RMSProp hiker keeps only a fading memory of recent squared slopes, so their caution tracks current conditions. The Adam hiker keeps RMSProp's fading memory, adds a fading memory of direction, and corrects both for the empty-notebook start.
| Method | Direction | Divides by | Bias correction | Relation to Adam |
|---|---|---|---|---|
| SGD | current gradient | nothing | n/a | no per-weight scaling at all |
| AdaGrad | current gradient | √(sum of all past g²) | n/a | Adam with β1 = 0, β2 → 1 and α decaying as 1/√t |
| RMSProp (with momentum) | momentum on the rescaled gradient | √(running average of g²) | no | closest relative; lacks the correction, which matters most when β2 is close to 1 |
| Adam | running average of g | √(running average of g²) | yes |
The paper also links Adam to natural gradient methods: dividing by √v̂ is a cheap, per-weight approximation to rescaling by the curvature of the problem, but a more conservative one.
Why it matters today
The family has kept growing (AdamW, Adafactor, Lion, Shampoo, Sophia), but nearly every member is still recognisably a running average of gradients divided by some estimate of their size.
6 Experiments · original
Everyday picture
The authors race optimizers on four courses of increasing difficulty and give every competitor its best settings, found by a grid search. Winning on a bowl-shaped course (logistic regression) is expected; the interesting courses are the bumpy ones (deep networks), where the theory says nothing.
| § | Problem | Setup | What the paper found |
|---|---|---|---|
| 6.1 | Logistic regression, MNIST digits | 784 pixel inputs, minibatch 128, step size decaying as α/√t | Adam converges about as fast as SGD with Nesterov momentum; both beat AdaGrad |
| 6.1 | Logistic regression, IMDB movie reviews | 10,000-word bag-of-words (very sparse) with 50% input dropout | Adam matches AdaGrad, which excels on sparse features, and beats SGD |
| 6.2 | Fully connected net, MNIST | 2 hidden layers of 1,000 ReLU units, minibatch 128, with and without dropout | Adam makes faster progress than the alternatives, including a quasi-Newton method (SFO) that was 5 to 10× slower per iteration |
| 6.3 | Convolutional net, CIFAR-10 | three 5×5 conv stages with pooling, then 1,000 ReLU units, minibatch 128 | Adam and SGD converge well ahead of AdaGrad; Adam's edge over tuned SGD is marginal but it adapts step sizes per layer automatically |
| 6.4 | Variational autoencoder, bias correction on and off | β1 ∈ [0, 0.9], β2 ∈ {0.99, 0.999, 0.9999}, α from 10−5 to 10−1 | Without correction, β2 near 1 causes instability early in training; with correction, Adam did as well as or better than RMSProp at every setting |
“In summary, Adam performed equal or better than RMSProp, regardless of hyper-parameter setting.”Kingma and Ba (2014), §6.4
Tiny example: why AdaGrad fades on the CNN
The paper notes that on the CNN, v̂ fell towards zero after a few epochs and was dominated by ε. AdaGrad's divisor, a sum of all past squared gradients, only grows: after 10,000 steps of gradients of size 0.1 it is √(10000 × 0.01) = 10, so every step is divided by 10 and keeps shrinking. A running average (RMSProp, Adam) of the same gradients stays at √0.01 = 0.1.
Why it matters today
“Works well with little tuning across very different models” is exactly the property that made Adam the default. The paper's own §6.3 also hints at the caveat still discussed today: carefully tuned SGD with momentum can match Adam on convolutional networks, and sometimes generalizes slightly better.
7 Extensions · original
7.1 AdaMax
Everyday picture
Adam's divisor is a fading average of squared gradients: a “typical size”. AdaMax replaces it with a fading record: the largest gradient seen recently, with old records slowly discounted. It is like judging how rough a road is by the worst pothole you remember, fading as it recedes behind you.
Tiny example
With β2 = 0.999, gradients of size 1 and then a single spike of 10: u jumps to max(0.999 × 1, 10) = 10 at the spike, then decays by 0.1% per step, 9.99, 9.98, …, so it takes about 2,300 steps to fall back to 1. Adam's √v̂ also jumps at the spike, but far less (to about 2.45 at step 20 in the chart below), and then eases back.
In words: “keep a slowly fading record of the largest gradient magnitude, and divide the bias-corrected average gradient by that record instead of by a root-mean-square.”
With the numbers: u = 1 before the spike; at the spike, max(0.999 × 1, |10|) = 10; next step, max(0.999 × 10, 1) = 9.99.
In Python:
beta2 = 0.999
# u before the spike
u = 1
# u_t = max(β2 · u_(t−1), |g_t|), with the spike g_t = 10
u = max(beta2 * u, abs(10))
u # → 10
# next step, the gradient is back to 1
u = max(beta2 * u, abs(1))
round(u, 2) # → 9.99
Hover or tap to compare the two divisors at any step.
Reading it: the gradient is 1 at every step except a spike of 10 at step 20. The AdaMax divisor u (orange, solid) leaps to the spike's size and then decays in a straight, slow line. The Adam divisor √v̂ (green, dashed) jumps only to about 2.45, because it averages the spike's square with everything before it, then eases back. The paper derives u as the limit of Adam's p-th power average as p → ∞; the max is what remains. AdaMax also needs no bias correction for u, because a max of values never drifts towards the made-up zero start.
7.2 Temporal averaging
Everyday picture
The last position of a jittery walker is a noisy guess of where they are heading; the average of their last few positions is steadier. Averaging the weights over the final stretch of training often generalizes better than taking the very last weights.
In words: “keep a fading average of the weights themselves, and use that average, corrected for its zero start by dividing by 1 − β2t, as the final model.”
With the numbers: with the weights 0.9, 0.8068, 0.7114 from §2 and β2 = 0.999, θ̄3 = 0.001 × (0.7114 + 0.999 × 0.8068 + 0.999² × 0.9) ≈ 0.002416, and dividing by 1 − 0.9993 = 0.002997 gives 0.806, close to the average of the three weights.
In Python:
beta2 = 0.999
# θ̄_0: the average starts at zero
theta_bar = 0
# θ_1, θ_2, θ_3 from §2
for theta in [0.9, 0.8068, 0.7114]:
# θ̄_t = β2 · θ̄_(t−1) + (1 − β2) θ_t
theta_bar = beta2 * theta_bar + (1 - beta2) * theta
round(theta_bar, 6) # → 0.002416
# the share of θ̄ that is real data
round(1 - beta2 ** 3, 6) # → 0.002997
# corrected for the zero start
round(theta_bar / (1 - beta2 ** 3), 3) # → 0.806
Why it matters today
Exponential moving averages of weights (often called EMA weights) are standard in image generation models and common in large-scale training. The Transformer paper used a close cousin: averaging the last few saved checkpoints.
8 Conclusion · original
“Overall, we found Adam to be robust and well-suited to a wide range of non-convex optimization problems in the field machine learning.”Kingma and Ba (2014), §8
The paper's summary is modest: AdaGrad's strength with sparse gradients plus RMSProp's strength with changing problems, cheap to compute, light on memory. History was less modest. Adam went on to become one of the most cited papers in computer science, and a nice detail from its footnotes survives it: the author order was decided by a coin flip.
What changed since 2014
| In the paper | Common today | Why | Where to learn more |
|---|---|---|---|
| L2 weight decay added to the gradient | AdamW: decay applied directly to the weights, separately from the adaptive step | Adding decay to the gradient lets √v̂ rescale it, so heavily updated weights are barely regularized (Loshchilov and Hutter, 2017) | optimizers lesson |
| β2 = 0.999 | β2 = 0.95 for large language models | A shorter memory for the squared gradient reacts faster to sudden gradient spikes, which destabilize very large training runs (GPT-3 used 0.95) | GPT-3 companion |
| α decaying as α/√t in the theory; fixed in most experiments | Warmup, then cosine decay | Warmup protects the fragile first steps; decay lets the model settle | optimizers lesson |
| Convergence proof for convex problems | Known to have a gap (Reddi et al., 2018; AMSGrad) | Rarely matters in practice, but a reminder that the empirical case is what carried Adam | |
| Two 32-bit numbers per weight of optimizer state | 8-bit optimizer states, sharded across GPUs | For billion-parameter models, the optimizer state outweighs the model itself | memory math in the inference lesson |
Glossary
Every term with hover guidance on this page, in one place.