Deep Residual Learning for Image Recognition, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it; so does every symbol in every equation.
- The diagrams are live: hover or tap any block to see what it does and the shape of the data there.
- The gradient-flow demo near the end lets you watch a signal fade through a deep plain network and survive through a residual one.
Each idea climbs the same ladder: an everyday picture, a tiny example, a diagram, the math, and why it matters today. Code that builds a residual connection and measures its effect lives in the deep networks lesson.
Abstract · original
“We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions.”He et al. (2015), Abstract
Everyday picture
Asking each layer of a network to produce a brand-new picture of the input is like asking every editor in a long chain to retype the whole manuscript from memory. Asking each editor only to mark corrections on the manuscript they were handed is far easier, and if a page is already fine they simply leave it alone. A residual network does the second: every block receives its input, computes a correction, and adds it to the input.
What the paper claims
- Deeper “plain” networks can be worse, even on the data they were trained on. Residual networks fix this.
- A 152-layer residual network, 8× deeper than the popular VGG nets, but cheaper to run.
- An ensemble reached 3.57% top-5 error on the ImageNet test set and won first place in the ILSVRC 2015 classification contest, plus first places in detection, localization and segmentation tracks.
- Networks of 100 and even 1,000 layers trained on CIFAR-10.
Why it matters today
The “+ x” this paper introduced is in every Transformer block, twice. Without residual connections there would be no hundred-layer language models.
1 Introduction · original
“Is learning better networks as easy as stacking more layers?”He et al. (2015), §1
Everyday picture
If a 20-storey building works, a 56-storey building made of the same floors should at least not be worse: you could always make the extra 36 floors empty corridors that pass everything straight through. The paper found that, in practice, deeper stacks of ordinary layers trained worse, even on their training data, so this was not overfitting. The training procedure simply could not find the “empty corridor” solution. The authors call this the degradation problem.
Tiny example
From the paper's Table 2 (ImageNet validation, top-1 error, lower is better): a plain 18-layer net scores 27.94% and a plain 34-layer net 28.54%, so 16 extra layers made it worse by 0.6 points. The residual versions, with the same layers plus shortcuts and no extra parameters, score 27.88% and 25.03%: the deeper one is now better by 2.85 points.
Reading it: each bar is one network's top-1 error on ImageNet validation images; shorter is better. Compare within each colour. Grey (plain): going from 18 to 34 layers makes the bar longer, which is the degradation problem. Blue (residual): the same step makes it clearly shorter. The 18-layer pair is nearly tied; depth is where shortcuts pay off. Numbers are selected from Table 2 of He et al. (2015).
The problem was not the classic vanishing gradient, which careful initialization and batch normalization had largely solved by 2015. The paper checks this: its plain networks use batch normalization and their gradients have healthy sizes. Something subtler was wrong.
Why it matters today
The lesson generalizes: “the model family can represent the solution” does not mean “the optimizer will find it”. How you parameterize a network changes what is easy to learn.
2 Related work · original
Everyday picture
Encoding differences instead of whole values is an old trick: video files store how each frame differs from the last, not every frame in full. The paper cites residual ideas from image retrieval (VLAD, Fisher vectors), from solving equations at multiple scales (multigrid methods), and older “shortcut connections” in neural networks.
The closest cousin
Highway networks, published months earlier, also had shortcuts, but controlled by learned gates that could close, and they had not shown gains beyond about 100 layers. ResNet's shortcut is gateless and parameter-free: the input always passes through.
Why it matters today
“Always on, no parameters” turned out to be the winning simplicity. Gated shortcuts survive in LSTMs, but the plain additive shortcut won in deep networks.
3 Deep residual learning · original
3.1 Residual learning · original
Everyday picture
A thermostat that is told “the right temperature is about where it already is, plus or minus a little” has an easy job. One told “work out the right temperature from scratch every time” has a hard one. If a stack of layers only has to learn the small difference between its input and the output it should produce, and that difference is often close to zero, learning gets easier.
Tiny example
Say the ideal output for an input x = 3 is H(x) = 3.2. A plain stack must build 3.2 out of nothing. A residual stack only has to learn F(x) = H(x) − x = 0.2. And if the ideal output is just x itself (the “empty corridor”), the residual stack gets it by setting its weights to zero, which is easy, while a plain stack of nonlinear layers must learn to imitate the identity, which turns out to be surprisingly hard.
In words: “instead of asking the layers to produce the desired output H(x) directly, ask them for the residual F(x), the desired output minus the input; the block then adds the input back.”
With the numbers: H(3) = 3.2, so F(3) = 3.2 − 3 = 0.2, and the block outputs F(3) + 3 = 3.2.
In Python:
# the input and the desired output H(x)
x, H_x = 3, 3.2
# F(x) := H(x) − x, what the layers learn
F_x = H_x - x
round(F_x, 1) # → 0.2
# the block adds x back: H(x) = F(x) + x
round(F_x + x, 1) # → 3.2
The paper tests the “residuals are small” hunch in §4.2: the layer outputs of residual networks are measurably smaller than those of plain networks, and shrink further as the network gets deeper. Each layer only nudges the signal.
Why it matters today
This framing, “each layer adds a small correction to a running representation”, is how researchers now describe Transformers too: the residual stream is carried through the whole model and every block writes a little into it.
3.2 Identity mapping by shortcuts · original
“Identity shortcut connections add neither extra parameter nor computational complexity.”He et al. (2015), §1
Everyday picture
A bypass road around a town: traffic can go through the town (the layers) or around it (the shortcut), and the two streams merge on the other side by simply adding up. The bypass costs nothing to maintain: it has no parameters and needs no computation beyond one addition.
Tiny example (2 numbers wide)
Input x = (1, 2). First weight layer W1 = [[0.5, 0], [0, −1]] gives (0.5, −2); ReLU keeps positives: (0.5, 0). Second weight layer W2 = [[0.2, 0], [0.4, 1]] gives F(x) = (0.1, 0.2). The shortcut adds the input back: y = (1, 2) + (0.1, 0.2) = (1.1, 2.2). The final ReLU leaves it unchanged. The block nudged its input by 10%.
Hover or tap a part. Start with x at the bottom.
Reading it: data flows upwards. The input x takes two routes. On the left it goes through two weight layers with a ReLU between them: that is F(x), the residual, inside the dashed box. On the right it skips everything along the identity shortcut. The two meet at ⊕, where they are added element by element, and a final ReLU follows. The only requirement is that F(x) and x have the same shape, so they can be added.
In words: “the block's output is the learned residual plus the input; when the residual has a different width than the input, the input is first passed through a small learned matrix Ws so the two can be added.”
With the numbers: x = (1, 2), W1x = (0.5, −2), σ(·) = (0.5, 0), W2(0.5, 0) = (0.1, 0.2), so y = (1.1, 2.2). For a stage that grows from 64 to 128 channels, Ws is a 1×1 convolution with 64 × 128 = 8,192 weights.
In Python:
# multiply a matrix by a vector, row by row
def matvec(W, v):
return [sum(w_j * v_j for w_j, v_j in zip(row, v)) for row in W]
x = [1, 2]
W1 = [[0.5, 0], [0, -1]]
W2 = [[0.2, 0], [0.4, 1]]
W1x = matvec(W1, x)
W1x # → [0.5, -2]
# σ keeps positives
relu = [max(0, v) for v in W1x]
relu # → [0.5, 0]
# F = W₂ σ(W₁ x)
F = matvec(W2, relu)
F # → [0.1, 0.2]
# y = F(x) + x
[F_k + x_k for F_k, x_k in zip(F, x)] # → [1.1, 2.2]
# weights in W_s, a 1×1 conv from 64 to 128 channels
1 * 1 * 64 * 128 # → 8192
Why it matters today
The shortcut's zero cost was essential to the paper's argument: plain and residual networks could be compared at exactly the same depth, width, parameter count and compute. That is why the comparison in §4 is so convincing.
3.3 Network architectures · original
Everyday picture
The network is a funnel of five stages. Each stage looks at a smaller, coarser image with more feature detectors: large and detailed with few detectors at the start, tiny and abstract with many at the end. Two design rules keep the cost per layer steady: when the image size halves, double the number of filters.
Try it: the five stages at each depth
Pick a depth above, then hover or tap a stage.
Reading it: the image enters on the left and shrinks at every stage (224 → 112 → 56 → 28 → 14 → 7 pixels across) while the number of filters grows (64 → 128 → 256 → 512). The “×n” under each stage is how many residual blocks it stacks; press the depth buttons and watch them change. Deeper networks add blocks mostly to the conv4_x stage (6 blocks at 34 layers, 23 at 101, 36 at 152). From 50 layers up, each block is a three-layer bottleneck (§4.1). At the end, average pooling squashes the 7×7 map to one vector and a single fully connected layer scores the 1,000 classes.
The paper's residual networks use identity shortcuts where shapes match. Where a stage changes width, it tries (A) padding the input with extra zeros (no parameters) or (B) a learned projection Ws.
Tiny example: cost
The 34-layer network needs 3.6 billion multiply-adds per image, only 18% of VGG-19's 19.6 billion (3.6 / 19.6 = 0.18), because it uses fewer filters and downsamples early.
Why it matters today
“Halve the resolution, double the channels, stack blocks, pool, classify” became the default template for image networks for years, and ResNet-50 became the standard benchmark model for measuring hardware speed.
3.4 Implementation · original
Everyday picture
A recipe card: nothing exotic. The point of listing it is that residual networks need no special training tricks.
- Data: images resized with a random shorter side between 256 and 480 pixels, a random 224×224 crop, random horizontal flips, standard colour augmentation.
- Normalization: batch normalization right after every convolution, before the activation.
- Optimizer: SGD with a batch of 256, momentum 0.9, weight decay 0.0001. The learning rate starts at 0.1 and is divided by 10 whenever the error stops improving, for up to 600,000 iterations.
- No dropout.
Why it matters today
The exact same optimizer and schedule train plain and residual networks. So when the residual networks win, the architecture is the only difference.
4 Experiments · original
4.1 ImageNet classification · original
Everyday picture
ImageNet is 1.28 million photos in 1,000 categories. “Top-1 error” is how often the model's single best guess is wrong; “top-5 error” is how often the right answer is not even among its five best guesses.
Three findings
- Residual learning reverses degradation. 34 residual layers beat 18 by 2.8 points and beat 34 plain layers by 3.5 points of top-1 error.
- Shortcut type barely matters. Zero-padding (A), projections only where widths change (B) and projections everywhere (C) give 25.03%, 24.52% and 24.19% top-1 error. All are far better than plain (28.54%), so projections are “not essential”, and the paper uses (B).
- Deeper keeps helping. With bottleneck blocks: 50 layers 22.85%, 101 layers 21.75%, 152 layers 21.43% top-1 (10-crop validation).
The bottleneck block
Everyday picture: to do expensive work on a wide road, first merge into fewer lanes, do the work, then fan back out. A 1×1 convolution cheaply squeezes 256 channels down to 64, the costly 3×3 convolution works on just 64, and another 1×1 restores 256.
Tiny example: the basic block on 64 channels has two 3×3 convolutions: 2 × (3 × 3 × 64 × 64) = 73,728 weights. The bottleneck on 256 channels has 1×1 (256 → 64) = 16,384, plus 3×3 (64 → 64) = 36,864, plus 1×1 (64 → 256) = 16,384, totalling 69,632. That is about the same cost, but it handles a representation four times wider.
Hover or tap a layer to see its weight count.
Reading it: both blocks read upwards and both have an identity shortcut on the right that jumps to the ⊕. On the left, two full-width 3×3 convolutions do the work. On the right, the purple 1×1 layers act as a funnel: squeeze 256 channels to 64, run the expensive 3×3 on the narrow version, then widen back to 256 so the shortcut's 256-d input can be added. The shortcut connects the two wide ends, which is why the paper insists it be a parameter-free identity: a projection there would double the block's cost.
In words: “a convolution layer has one weight for every position in its k × k window, for every input channel, for every output filter.”
With the numbers: 3 × 3 × 64 × 64 = 36,864; 1 × 1 × 256 × 64 = 16,384.
In Python:
k, C_in, C_out = 3, 64, 64
k * k * C_in * C_out # → 36864
k, C_in, C_out = 1, 256, 64
k * k * C_in * C_out # → 16384
How deep versus how good
Reading it: each bar is one network's top-1 error on ImageNet validation images (10-crop testing); shorter is better. VGG-16 and the 34-layer plain net sit at the top at about 28%. Adding shortcuts to the 34-layer net drops it to about 25%, and every further increase in depth, 50, 101, then 152 layers, shortens the bar again. There is no sign of degradation. Numbers selected from Table 3 of He et al. (2015).
The 152-layer network's single-model top-5 validation error was 4.49%, better than every previous ensemble. An ensemble of six networks of different depths reached 3.57% top-5 on the test set and won ILSVRC 2015. It is also cheaper than VGG: 11.3 billion FLOPs against VGG-16's 15.3 billion.
Why it matters today
Depth had been the axis everyone wanted to push and couldn't. After this paper, depth was cheap, and the question became how to use it.
4.2 CIFAR-10 and analysis · original
Everyday picture
CIFAR-10 is 60,000 tiny 32×32 photos in 10 classes: a laboratory where you can afford to train absurdly deep networks just to see what happens. The authors use a deliberately simple family: three stages of 2n layers each, plus a first and a last layer, 6n + 2 layers in total.
Tiny example
n = 3 gives 6 × 3 + 2 = 20 layers; n = 18 gives 110 layers; n = 200 gives 1,202 layers. The plain versions degrade (the paper reports the 110-layer plain net above 60% error), while the residual versions keep improving up to 110 layers.
Hover or tap to read the error at each depth.
Reading it: the x-axis is depth on a log scale; the y-axis is test error on CIFAR-10. The residual networks improve steadily from 20 layers (8.75%) to 110 layers (6.43%). Then the 1,202-layer network gets worse (7.93%), even though it trains to a similar training error. The paper attributes that to overfitting: 19.4 million parameters on only 50,000 training images, with no dropout. So the ceiling moved from “can't be optimized” to ordinary overfitting. Numbers selected from Table 6 of He et al. (2015).
Small residuals, measured
The paper measures how strongly each layer responds (the standard deviation of its output). Residual networks respond more weakly than plain ones, and the deeper the residual network, the weaker each layer. That supports §3.1's hunch: when layers only add corrections, each correction is small.
Why it matters today
The 110-layer CIFAR ResNet trained with a brief warmup (learning rate 0.01 until training error fell below 80%, about 400 iterations, then 0.1): an early sighting of the warmup trick that is now standard.
4.3 Object detection on PASCAL and MS COCO · original
Everyday picture
If a better backbone learns better features, it should improve every task built on it. The paper swaps VGG-16 for ResNet-101 inside the same Faster R-CNN detector and changes nothing else.
Tiny example
On COCO's main metric, mAP averaged over overlap thresholds from 0.5 to 0.95, the score rises from 21.2 to 27.2: +6.0 points, and 6.0 / 21.2 = a 28% relative improvement, from the backbone alone.
Why it matters today
“Pretrain a deep residual backbone, then reuse it for everything” is the pattern that foundation models later scaled up.
Why gradients like shortcuts
This section is our addition. The paper's own argument (§4.1) is about optimization: with batch normalization its plain networks' gradients were healthy, yet they still trained worse. A follow-up paper by the same authors (He et al., 2016, “Identity Mappings in Deep Residual Networks”) made the gradient view precise, and it is the view most people carry today.
Everyday picture
Training sends a message (the gradient) backwards from the output to every layer. In a plain network the message is retold by every layer in turn, and a message retold forty times by people who each drop a little of it arrives as a whisper. A residual network adds an express lane: the ⊕ passes the message straight through every block, untouched, alongside whatever the block adds.
In words: “how much the block's output moves when its input moves is the identity (a copy) plus whatever the residual branch contributes; even if the branch contributes almost nothing, the copy is still there.”
With the numbers: for a one-number block y = x + 0.1x, the slope is 1 + 0.1 = 1.1. Stack 40 plain layers that each shrink the signal to 0.5 of its size and the gradient reaching the first layer is 0.540 ≈ 10−12; with shortcuts each factor is 1 + (something small) instead.
In Python:
# ∂F/∂x for the block y = x + 0.1x
dF_dx = 0.1
# ∂y/∂x = I + ∂F/∂x
1 + dF_dx # → 1.1
# 40 plain layers, each halving: about 10⁻¹²
0.5 ** 40 # → 9.09...e-13
Hover or tap to read the gradient size at any depth.
Reading it: this is a live simulation, not a result from the paper: a 40-layer network, 16 units wide, with random weights and no batch normalization. The x-axis is the layer, from the input (0) to the output (40). The y-axis is how large the gradient is at that layer, relative to the output, on a log scale: 0 means “the same size”, −6 means “a millionth”. Start with weight scale at 0.5, small initial weights. The plain network (blue, solid) shrinks the gradient by a factor of about 1013 by the time it reaches the first layer: its early layers cannot learn. The residual network (green, dashed) never falls below about the output's size, because every ⊕ passes the gradient through. Now raise the weight scale to 1 (the He initialization): the plain network now stays within about a factor of ten, but the residual one grows, by a factor of several thousand, backwards through the blocks, since every block adds its own contribution on top of the copy. Finally lower branch scale, which multiplies each residual branch by a small number: the green dashed line flattens near 0, a network that starts out as nearly the identity. That is why the paper pairs shortcuts with batch normalization, and why modern practice often initializes the last layer of every residual branch at or near zero.
Why it matters today
Pre-norm Transformers keep a clean identity path from the last layer all the way to the first (normalization sits only inside the branches), for exactly this reason. See the deep networks lesson, which measures gradient flow through plain and residual stacks in NumPy.
What changed since 2015
| In the paper | Common today | Why | Where to learn more |
|---|---|---|---|
| ReLU after the addition | “Pre-activation”: normalization and ReLU inside the branch, nothing after the add | Keeps a perfectly clean identity path end to end (He et al., 2016) | deep nets |
| Convolutional blocks | Attention and feed-forward blocks in Transformers, each wrapped in the same x + F(x) | The shortcut is architecture-agnostic | Transformer companion |
| Batch normalization | Layer normalization or RMSNorm in language models | Batch statistics are awkward for variable-length sequences and small batches | Layer Norm companion |
| CNNs as the vision workhorse | Vision Transformers at large scale; ResNets still common when data or compute is limited | Attention scales better with data; convolutions' built-in assumptions help on small data | CNNs and RNNs |
Glossary
Every term with hover guidance on this page, in one place.