An annotated companion · AI Primer

Attention Is All You Need, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The equations and the paper's tables are reproduced with attribution under the notice printed on the paper itself: “Provided proper attribution is provided, Google hereby grants permission to reproduce the tables and figures in this paper solely for use in journalistic or scholarly works.” The architecture figures here are redrawn from scratch. Read the original alongside: every section links to it.

How to read this page

Nothing on this page assumes you already know the jargon. Three things help:

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same. Hover the Q in the first equation below and see.
  • The diagrams are live: hover or tap any block to see what it does and what shape the data has at that point.

Each idea is explained in the same order: an everyday picture, a tiny example you could check by hand, a diagram, then the math, and finally why it still matters. Code that builds each piece from scratch lives in the linked lessons (for example the attention lesson).

Abstract

“We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.”Vaswani et al. (2017), Abstract. Read the original

Everyday picture

Before 2017, the best translation systems read a sentence the way you'd read through a keyhole: one word at a time, carrying a running summary in your head. This paper says: lay the whole sentence on the table, and let every word look at every other word directly. That “looking” is attention, and a model built only from attention (plus some simple per-word processing) is the Transformer.

What the paper claims

  • A translation model with no recurrence and no convolutions: attention only.
  • Better quality: 28.4 BLEU on English→German and 41.8 on English→French, new records at the time.
  • Much cheaper to train, because every word can be processed in parallel: the big model trained in 3.5 days on 8 GPUs.

Why it matters today

Almost every modern language model (GPT, Claude, Llama, Gemini), most image and speech models, and every modern embedding model is a descendant of the architecture in this paper. Understanding these 15 pages is understanding the core of modern AI.

1 Introduction · original

“This inherently sequential nature precludes parallelization within training examples, which becomes critical at longer sequence lengths, as memory constraints limit batching across examples.”Vaswani et al. (2017), §1

Everyday picture

Imagine a relay race where each runner can only start when the previous one hands over the baton. Adding more runners doesn't make the race faster. That is a recurrent neural network (RNN): to process word 50, it must first finish words 1 to 49, because each step needs the hidden state from the step before. A GPU, which is built to do thousands of things at once, sits mostly idle.

Tiny example

Take “The cat sat down”. An RNN computes four steps in a strict chain: h1 from “The”; h2 from h1 and “cat”; h3 from h2 and “sat”; h4 from h3 and “down”. Information about “The” reaches “down” only after surviving three hand-offs. The Transformer computes all four words' new representations at once, and “down” can look at “The” in a single step.

Why it matters

This is the paper's core bet: give up the step-by-step memory, and gain parallelism plus direct connections. The bet paid off so well that training on vastly more data became practical, which is what eventually produced today's large language models. The CNN and RNN lesson builds an RNN from scratch and measures exactly how it forgets.

2 Background · original

Everyday picture

Others had already tried to escape the relay race. Convolution-based models (ByteNet, ConvS2S) read words in small overlapping windows, like reading a page through a sliding frame a few words wide. Windows run in parallel, but for two far-apart words to “meet”, the frames have to be stacked many layers deep.

The key idea the paper borrows

Self-attention (relating different positions of one sequence to each other) already existed as a helper inside recurrent models. The new move is making it the only mechanism. The paper also flags the price: averaging over many positions blurs detail, which is why it introduces multi-head attention (§3.2.2): several attentions side by side, each free to focus on something different.

Why it matters

“Any two words are one step apart” is the property that makes Transformers good at long-range relationships, such as which noun a pronoun refers to 30 words later.

3 Model Architecture · original

Everyday picture

Think of a two-person translation team. The encoder reads the whole English sentence and writes rich notes about every word. The decoder writes the German sentence one word at a time; before each word, it rereads what it has written so far and consults the encoder's notes. Writing one word at a time, each time feeding the output back in, is what auto-regressive means.

Hover or tap any block in the redrawn figure to see what it does and the shape of the data flowing through it. The numbers use the paper's base model: dmodel = 512, h = 8 heads, N = 6 layers, and a 10-token sentence as the example.

Encoder Decoder N× N× Inputs Input Embedding +PositionalEncoding Multi-HeadAttention Add & Norm Feed Forward Add & Norm Outputs (shifted right) Output Embedding +PositionalEncoding Masked Multi-HeadAttention Add & Norm Multi-HeadAttention Add & Norm Feed Forward Add & Norm Linear Softmax Output Probabilities encoder output (becomes K, V)

Hover or tap any block in the diagram. Start at the bottom left with Inputs and work upwards.

Figure 1 of the paper, redrawn: the Transformer's encoder (left) and decoder (right). Based on Vaswani et al. (2017), Figure 1, reproduced in redrawn form with attribution.

Reading it: data flows from the bottom up. On the left, the input sentence becomes vectors (Input Embedding), gets word-order information added (the ⊕ with the wave), and passes through the encoder block, which is repeated N = 6 times (the dashed box). Each repetition has two steps: attention, where words look at each other, and a feed-forward network, where each word is processed on its own. Each step is followed by Add & Norm. The encoder's final output travels right along the long wire and feeds the decoder's middle attention block. On the right, the decoder does the same climb over the words written so far, with one extra attention step in the middle to consult the encoder. At the top, Linear and Softmax turn the final vector into a probability for every word in the vocabulary. The arrows that bypass a block and rejoin at Add & Norm are the residual connections.

3.1 Encoder and decoder stacks · original

Everyday picture

Each layer is a draft that the next layer improves. Instead of rewriting the draft from scratch, each layer writes corrections in the margin that get added to the draft: that is the residual connection. Then the page is tidied to a standard size so no number grows out of control: that is layer normalization.

Tiny example

Say a word's vector is x = (2, 0) and the attention step proposes the change Sublayer(x) = (1, 1). The residual adds them: (3, 1). Layer norm then rescales that to zero mean and unit spread: the mean is 2 and the standard deviation is 1, so the result is ((3 − 2)/1, (1 − 2)/1) = (1, −1).

In words: “add the sub-layer's proposed change to the input, then normalize the sum.”

With the numbers: LayerNorm((2, 0) + (1, 1)) = LayerNorm((3, 1)) = (1, −1).

In Python:

import math
x = [2, 0]
# Sublayer(x): the proposed change
sublayer_x = [1, 1]
# x + Sublayer(x): the residual
s = [x_i + c_i for x_i, c_i in zip(x, sublayer_x)]
s  # → [3, 1]
mean = sum(s) / len(s)
std = math.sqrt(sum((s_i - mean) ** 2 for s_i in s) / len(s))
# LayerNorm: zero mean, unit spread
[(s_i - mean) / std for s_i in s]  # → [1.0, -1.0]

What the paper specifies

  • Encoder: N = 6 identical layers. Each has two sub-layers, multi-head self-attention and a feed-forward network, each wrapped as above.
  • Decoder: N = 6 identical layers with a third sub-layer in the middle that attends over the encoder's output. Its self-attention is masked so position i can only see positions before it.
  • Every sub-layer and embedding outputs vectors of width dmodel = 512, so the pieces snap together like Lego of one fixed size.

Why it matters today

Residual connections are why you can stack 6 layers here and over a hundred in modern models: gradients flow straight down the “add” path. One detail has changed: modern models normalize before each sub-layer (x + Sublayer(LayerNorm(x)), called pre-norm), which trains more stably when very deep. See the transformer lesson and the deep networks lesson.

3.2 Attention · original

“An attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, keys, values, and output are all vectors.”Vaswani et al. (2017), §3.2

Everyday picture

You walk into a library and type a search (your query). Every book has a catalogue card describing it (its key) and the book itself on the shelf (its value). The librarian compares your search with every card and hands you not one book but a blend: mostly the best match, a little of the next best. In a Transformer, every word is simultaneously a searcher, a catalogue card and a book: this is self-attention.

Try it: who does each word listen to?

Last word of the sentence:

Hover, tab to or tap a word. The bars show how much attention it pays to each word.

These weights are illustrative, chosen by hand to show the kind of pattern trained models learn. They are not taken from a real model.

Reading it: pick a word and look at the bars under every word: the longer the bar, the more of that word's meaning flows into the one you picked. Pick it. With “tired”, most attention goes to animal, because animals get tired. Switch the ending to “wide” and it now attends mostly to street, because streets are wide. One changed word at the end changes what “it” means, and attention is the mechanism that carries that change. No grammar rule does this; it is learned from data.

Why it matters today

Attention is the only place in a Transformer where words exchange information. Everything else processes each word on its own. If you understand this section, you understand the part of the model that makes it contextual.

3.2.1 Scaled dot-product attention · original

Tiny example

Follow the word “it”. Its query is compared with three keys using a dot product, and the results are scaled. Say the scaled scores come out as 2.0 for “animal”, 1.0 for “tired” and 0.5 for “street”. Softmax turns them into shares: e2.0 = 7.39, e1.0 = 2.72 and e0.5 = 1.65, which total 11.76. So “animal” gets 7.39 / 11.76 = 0.63, “tired” gets 0.23 and “street” gets 0.14. The new vector for “it” is 0.63 × value(animal) + 0.23 × value(tired) + 0.14 × value(street).

Q K V MatMul Scale Mask (optional) SoftMax MatMul Output

Hover or tap a step. Start with Q and K at the bottom.

Figure 2 (left) of the paper, redrawn: scaled dot-product attention. Based on Vaswani et al. (2017), Figure 2.

Reading it: read from the bottom up. Q and K enter the first MatMul, which produces the grid of scores: every query against every key. Scale divides by √dk. The optional Mask blanks out positions a word is not allowed to see. SoftMax turns each row of scores into shares that sum to 1. V has been waiting on the right all along: the second MatMul uses the shares to blend the values. Q and K decide how much; V carries what.

The math

In words: “score every query against every key, shrink the scores by the square root of the key width, turn each row of scores into shares, and use the shares to blend the values.”

With the numbers: the “it” row of QKᵀ/√dk is (2.0, 1.0, 0.5); softmax makes it (0.63, 0.23, 0.14); multiplying by V gives 0.63·V(animal) + 0.23·V(tired) + 0.14·V(street).

In Python:

import math
# the "it" row of Q Kᵀ / √d_k
scores = [2.0, 1.0, 0.5]
total = sum(math.exp(z) for z in scores)
# softmax: one share per word
weights = [math.exp(z) / total for z in scores]
[round(w, 2) for w in weights]  # → [0.63, 0.23, 0.14]
# stand-in values: animal, tired, street, one direction each
V = [[1, 0, 0], [0, 1, 0], [0, 0, 1]]
# shares times V
[round(sum(w * v[d] for w, v in zip(weights, V)), 2) for d in range(3)]  # → [0.63, 0.23, 0.14]

Softmax, decoded

Softmax is the step that turns any list of scores, including negative ones, into shares that are all positive and add up to exactly 1. Think of it as a vote in which louder voices get disproportionately more say.

In words: “the share for item i is e raised to its score, divided by the sum of e raised to every score.”

With the numbers: softmax(2.0, 1.0, 0.5) for “animal” = 7.39 / (7.39 + 2.72 + 1.65) = 7.39 / 11.76 = 0.63.

In Python:

import math
# scores for animal, tired, street
z = [2.0, 1.0, 0.5]
# e^(z_j) for each j
[round(math.exp(z_j), 2) for z_j in z]  # → [7.39, 2.72, 1.65]
# Σ_j e^(z_j)
denominator = sum(math.exp(z_j) for z_j in z)
round(denominator, 2)  # → 11.76
# softmax(z)_i for i = animal
round(math.exp(z[0]) / denominator, 2)  # → 0.63

Reading it: drag a score and watch every bar move, because the shares always add up to 1: one word can only gain attention by taking it from the others. Now drag temperature. Softmax is computed on score ÷ temperature, so a low temperature exaggerates the gaps (the winner takes nearly everything) and a high temperature flattens them (everyone gets a similar share). The same knob controls how adventurous a chat model is when it picks its next word; see the inference lesson.

Why divide by √dk? The paper's footnote, decoded

“We suspect that for large values of dk, the dot products grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients.”Vaswani et al. (2017), §3.2.1

Everyday picture: roll one die and the result swings between 1 and 6. Add up a hundred dice and the total swings by dozens. A dot product is a sum of dk small products, so the wider the vectors, the wilder the scores. Wild scores make softmax give one word nearly 100%, and then it stops responding to small changes, so the gradient that training relies on disappears.

In words: “a dot product adds up dk products; if each entry of q and k is random with average 0 and spread 1, the sum's variance is dk, so its typical size is √dk. Dividing by √dk brings the typical size back to 1.”

With the numbers: the paper's dk = 64, so raw scores typically swing by √64 = 8. After dividing by 8 they typically swing by about 1, where softmax is smooth.

In Python:

import math
d_k = 64
# Var(q·k) = d_k: each of the d_k products adds spread 1
variance = d_k
# typical swing of a raw score
math.sqrt(variance)  # → 8.0
# after dividing by √d_k
math.sqrt(variance) / math.sqrt(d_k)  # → 1.0

Why it matters today

Every attention layer in every modern model still divides by √dk. The attention lesson measures the saturation directly: at dk = 512 without scaling, the largest share averages above 0.9 and the gradient collapses.

3.2.2 Multi-head attention · original

“Multi-head attention allows the model to jointly attend to information from different representation subspaces at different positions.”Vaswani et al. (2017), §3.2.2

Everyday picture

Give the same sentence to eight readers, each with a different coloured highlighter. One marks who did what, another marks which pronoun points where, another marks nearby words. Afterwards you staple their notes together. That is multi-head attention: eight small attentions side by side, each free to look for something different.

Tiny example

The base model's vectors have 512 numbers. Instead of one attention over all 512, it runs h = 8 heads over 512 / 8 = 64 numbers each. Each head has its own projection matrices of size 512 × 64. The eight 64-wide outputs are glued back into 8 × 64 = 512 numbers and mixed by one more 512 × 512 matrix. The total work is about the same as one full-width head.

V K Q Linear Linear Linear Scaled Dot-ProductAttention h Concat Linear Output

Hover or tap a part. The stacked shadows mean “h copies of this, side by side”.

Figure 2 (right) of the paper, redrawn: multi-head attention with h parallel heads. Based on Vaswani et al. (2017), Figure 2.

Reading it: the shadows behind a box mean “h copies of this, running side by side”. V, K and Q each pass through their own small Linear projection once per head, which cuts them down to 64 numbers. Each head then runs the whole recipe from the previous figure independently. Concat glues the h results back into one 512-wide vector per word, and the final Linear (WO) lets the heads' findings mix.

The math

In words: “for each head i, project the queries, keys and values down with that head's own matrices and run attention; then glue all the heads' outputs side by side and mix them with one more matrix.”

With the numbers: a 10-word sentence gives Q, K, V of shape 10 × 512. Each WiQ is 512 × 64, so each head sees 10 × 64. Eight heads concatenated give 10 × 512, and WO (512 × 512) keeps it at 10 × 512. The layer's projection weights total 4 × 512 × 512 = 1,048,576 numbers.

In Python:

n, d_model, h = 10, 512, 8
# each W_i^Q is d_model × d_k
d_k = d_model // h
# shape of Q W_i^Q: what one head sees
(n, d_k)  # → (10, 64)
# Concat(head_1, ..., head_h)
(n, h * d_k)  # → (10, 512)
# W^Q, W^K, W^V (8 heads each) and W^O
print(f"{4 * d_model * d_model:,}")  # → 1,048,576

Why it matters today

Researchers have found heads that specialise: some track the previous word, some copy patterns, some resolve pronouns. Modern models keep many query heads but share keys and values across groups of heads (grouped-query attention), which shrinks the memory needed while generating.

3.2.3 Three uses of attention in the model · original

Everyday picture

The same “look around and blend” tool is used in three places, differing only in who asks and who answers: the reader re-reading their own notes (encoder self-attention), the writer re-reading their own draft without peeking at words not yet written (masked decoder self-attention), and the writer consulting the reader's notes (encoder-decoder attention).

WhereQueries come fromKeys and values come fromMask
Encoder self-attentionthe input sentencethe same input sentencenone: every word sees every word
Decoder self-attentionthe output written so farthe same outputcausal: no looking ahead
Encoder-decoder attentionthe decoderthe encoder's final outputnone

Tiny example: the causal mask

keys (looked at) →

Hover or tap a cell: row = the word looking, column = the word looked at.

The causal mask for the four-word output “Die Katze saß .”. Blue cells are allowed; grey cells are set to −∞ before softmax.

Reading it: each row is one word looking back at the sentence. Blue means allowed and grey means masked. The allowed cells form a lower triangle: the first word can see only itself, and the last sees everything before it. The masked scores are set to −∞ before softmax, and e−∞ = 0, so their shares are exactly zero. During training, this is what stops the decoder from cheating by reading the word it is supposed to predict.

Why it matters today

GPT, Claude and Llama keep only the middle row of that table: they are decoder-only models, a single stack of masked self-attention. Encoder-only models (BERT) keep only the first row and power most embedding models. The transformer lesson compares the three families.

3.3 Position-wise feed-forward networks · original

Everyday picture

Attention is the meeting where everyone listens to everyone. The feed-forward network is what happens afterwards: each person goes back to their desk and thinks on their own about what they heard. “Position-wise” means the same desk-work recipe is applied to every word separately.

Tiny example (2 numbers wide instead of 512)

Take a word vector x = (1, −1). First, multiply by W1 = [[1, 0, −1], [0, 1, 1]] to widen it to 3 numbers: (1·1 + (−1)·0, 1·0 + (−1)·1, 1·(−1) + (−1)·1) = (1, −1, −2). Next, apply ReLU, max(0, ·), which keeps positives and zeroes negatives: (1, 0, 0). Finally, multiply by W2 = [[2, 0], [0, 1], [1, 1]] to come back to 2 numbers: (2, 0). The biases are zero here to keep it short.

In words: “widen the word's vector by a factor of four, zero out the negatives, and narrow it back to its original width.”

With the numbers: in the base model, x has 512 numbers, W1 widens to dff = 2048, and W2 returns to 512. That is 2 × 512 × 2048 + 2048 + 512 = 2,099,712 parameters per layer, about twice the attention layer's projections.

In Python:

# the tiny example, 2 wide
x = [1, -1]
W1 = [[1, 0, -1], [0, 1, 1]]
W2 = [[2, 0], [0, 1], [1, 1]]
# x W1 (b1 = 0)
xW1 = [sum(x[r] * W1[r][c] for r in range(2)) for c in range(3)]
xW1  # → [1, -1, -2]
# max(0, ...): ReLU
hidden = [max(0, a) for a in xW1]
# ... W2 (b2 = 0)
[sum(hidden[r] * W2[r][c] for r in range(3)) for c in range(2)]  # → [2, 0]
d_model, d_ff = 512, 2048
# W1, b1, W2, b2
print(f"{d_model * d_ff + d_ff + d_ff * d_model + d_model:,}")  # → 2,099,712

Why it matters today

Feed-forward layers hold about two-thirds of a Transformer's parameters and are thought to store much of its factual knowledge. Modern models swap ReLU for smoother functions (GELU, SwiGLU), and mixture-of-experts models replace this one network with many, routing each word to a few.

3.4 Embeddings and softmax · original

Everyday picture

An embedding table is a phone book with one row per token in the vocabulary. Looking a word up returns its 512 numbers. At the other end, the model has to turn its final 512 numbers back into a choice of word. The paper uses the same phone book backwards: it compares the final vector with every row and asks which row it most resembles.

Tiny example

If “cat” is token 17, its embedding is simply row 17 of the table, multiplied by √dmodel = √512 ≈ 22.6. The paper scales it up so the word's meaning isn't drowned out by the positional encoding that gets added next. At the output, a final vector h is dotted with every row, giving one score (a logit) per vocabulary entry. Softmax turns those scores into probabilities for the next word.

In words: “score the final vector against every word's embedding, then turn the scores into probabilities.”

With the numbers: with a vocabulary of about 37,000 tokens, E is 37,000 × 512 = 18,944,000 numbers. Sharing it between the input, the output and the final linear layer saves two more copies of that.

In Python:

K, d_model = 37_000, 512
# numbers in E: one row of d_model per token
print(f"{K * d_model:,}")  # → 18,944,000

Why it matters today

Every language model ends exactly like this: a score per vocabulary token, then softmax, then sampling. How text is split into tokens in the first place is covered in the tokenization lesson.

3.5 Positional encoding · original

“We chose this function because we hypothesized it would allow the model to easily learn to attend by relative positions, since for any fixed offset k, PEpos+k can be represented as a linear function of PEpos.”Vaswani et al. (2017), §3.5

Everyday picture

Attention on its own has no sense of order: “dog bites man” and “man bites dog” contain the same words, so they get the same result. The fix is to stamp each word with its position. The paper's stamp is like a wall of clocks whose hands turn at different speeds: a seconds hand, a minutes hand, an hours hand, and on up to hands that barely move. Reading every hand together tells you the exact time, and the gap between two times is always the same set of hand rotations, which makes “three words later” easy to recognise.

Tiny example

For position 1, the fastest pair of dimensions (i = 0) gets sin(1) = 0.841 and cos(1) = 0.540. The next pair (i = 1) turns slightly slower: sin(1 / 1.0366) = 0.822 and cos = 0.570. At position 10, dimension pair i = 64 divides by exactly 10, so it shows sin(1) = 0.841 again. Every position gets a unique 512-number stamp, and it is added to the word's embedding.

In words: “for position pos, each pair of dimensions gets a sine and a cosine of the position, divided by a number that grows from 1 to 10,000 as you move across the pairs, so early pairs spin fast and late pairs spin slowly.”

With the numbers: PE(1, 0) = sin(1/1) = 0.841; PE(1, 1) = cos(1/1) = 0.540; PE(10, 128) = sin(10/10) = 0.841, because pair i = 64 has 10000128/512 = 10.

In Python:

import math
d_model = 512
def PE(pos, dim):
    # the pair this dimension belongs to
    i = dim // 2
    angle = pos / 10000 ** (2 * i / d_model)
    # even: sin, odd: cos
    return math.sin(angle) if dim % 2 == 0 else math.cos(angle)
print(f"{PE(1, 0):.3f} {PE(1, 1):.3f}")  # → 0.841 0.540
# pair i = 64 divides by exactly 10
10000 ** (128 / 512)  # → 10.0
print(f"{PE(10, 128):.3f}")  # → 0.841
Positions 0–99 (rows) × the first 128 dimensions (columns), computed live from the formula with dmodel = 512.

Hover or tap the heatmap to read a value.

−10+1

Reading it: each row is one position's stamp; each column is one dimension. On the left, the fast-spinning dimensions flip between red (+1) and blue (−1) every few rows. Moving right, the stripes stretch out, because those dimensions spin more slowly. Far right, they barely change over 100 positions. Read any row across and you get a barcode unique to that position. Hover to see the exact position, dimension and value.

Why it matters today

The paper also tried learned position vectors and got nearly identical results (Table 3, row E). Most modern models instead use RoPE, which rotates queries and keys by an angle set by position, using the same clock-hand idea. See the positional encoding lesson and the RoFormer companion.

4 Why self-attention · original

Everyday picture

How many phone calls does it take for news to travel between two people? In a chain of friends (an RNN), it takes as many calls as there are people between them. In a group call (self-attention), it takes one. The paper compares layer types on three questions: how much work per layer, how much must happen in sequence, and how many hops information needs between any two words.

Table 1, reproduced with attribution (Vaswani et al., 2017)
Layer typeWork per layerSequential stepsLongest path between two words
Self-attentionO(n²·d)O(1)O(1)
RecurrentO(n·d²)O(n)O(n)
ConvolutionalO(k·n·d²)O(1)O(logk(n))
Self-attention (restricted to a window of r)O(r·n·d)O(1)O(n/r)

Here n is the sentence length, d the vector width, k the convolution window and r the attention window. O(·) describes how cost grows, ignoring constant factors.

Tiny example: when is n² cheaper than d²?

d = 512

Reading it: the two bars compare the work per layer of self-attention (n²·d) and of a recurrent layer (n·d²). Slide n. While sentences are shorter than the vector width (n < d = 512), self-attention is cheaper, and in 2017 most sentences were a few dozen tokens long. Past n = 512 the balance flips. Today's contexts run to hundreds of thousands of tokens, which is why so much modern engineering (FlashAttention, sliding windows, the KV cache) goes into taming the n² term.

Why it matters today

The two columns that won the argument are the last two: O(1) sequential steps means training uses the whole GPU, and an O(1) path means distant words connect directly. The paper's closing aside about attention being more interpretable, since you can inspect what attends to what, started a whole research field.

5 Training · original

5.1–5.2 Data, hardware and schedule

  • Data: the WMT 2014 English→German set (about 4.5 million sentence pairs, split into a shared vocabulary of about 37,000 sub-word tokens with byte-pair encoding) and English→French (36 million sentences, 32,000 word-piece tokens).
  • Batches: sentences of similar length grouped together, about 25,000 source and 25,000 target tokens per batch.
  • Hardware: one machine with 8 NVIDIA P100 GPUs. The base model trained for 100,000 steps (about 12 hours, 0.4 s per step); the big model for 300,000 steps (3.5 days, 1.0 s per step).

Everyday picture: twelve hours on one machine was remarkable in 2017. Competing systems needed several times more compute for worse results, and that efficiency is the whole case for giving up recurrence.

5.3 Optimizer and the warmup schedule · original

Everyday picture

A new driver pulls out of the driveway slowly, speeds up on the open road, then eases off as they near the destination. The learning rate (how big each training step is) follows the same shape. It starts tiny, because the untrained model's gradients point in wild directions. It climbs during warmup, then decays so the model can settle into fine detail. The optimizer is Adam with β1 = 0.9, β2 = 0.98 and ε = 10−9.

In words: “for the first warmup_steps steps, grow the learning rate in a straight line; after that, shrink it in proportion to one over the square root of the step number; scale everything by one over the square root of the model width.”

With the numbers: with dmodel = 512 and warmup_steps = 4000, the peak is at step 4000: 512−0.5 × 4000−0.5 = 0.0442 × 0.0158 ≈ 7.0 × 10−4. By step 100,000 it has decayed to 512−0.5 × 100000−0.5 ≈ 1.4 × 10−4.

In Python:

d_model, warmup_steps = 512, 4000
def lrate(step_num):
    return d_model ** -0.5 * min(step_num ** -0.5, step_num * warmup_steps ** -1.5)
round(d_model ** -0.5, 4), round(4000 ** -0.5, 4)  # → (0.0442, 0.0158)
# the peak, in units of 10^-4
round(lrate(4000) * 1e4, 1)  # → 7.0
# decayed by step 100,000
round(lrate(100_000) * 1e4, 1)  # → 1.4

Hover or tap the curve to read the learning rate at any step.

Reading it: the x-axis is the training step and the y-axis the learning rate. The curve rises in a straight line to a peak exactly at warmup_steps, then falls away along a 1/√step curve. Drag the slider: a longer warmup gives a lower, later peak, and a shorter warmup a higher, earlier one. After the peak, every setting follows the same decay curve, because the warmup term no longer matters.

Why it matters today

Warmup followed by decay is still standard for training Transformers; today the decay is usually a cosine curve, and the optimizer is usually AdamW. The optimizers lesson shows what happens without warmup.

5.4 Regularization · original

Everyday picture

Two tricks keep the model from memorising (overfitting). Dropout is like a sports team that trains with random players benched each practice, so nobody can rely on one star. Label smoothing is like a teacher who marks the right answer as “90% right” instead of “100% right”, so students never learn to be absolutely certain.

  • Residual dropout, Pdrop = 0.1: 10% of each sub-layer's outputs are zeroed at random before the Add & Norm, and the same happens to the sums of embeddings and positional encodings.
  • Label smoothing, εls = 0.1: the target for the correct token becomes 0.9 plus a tiny share, and the remaining 0.1 is spread over all tokens.

In words: “take 90% of the true one-hot target and add an equal sliver of the remaining 10% to every token in the vocabulary.”

With the numbers: with K = 37,000 and εls = 0.1, the correct token's target is 0.9 + 0.1 / 37,000 ≈ 0.900003, and every other token's is 0.1 / 37,000 ≈ 0.0000027.

In Python:

eps_ls, K = 0.1, 37_000
def smooth(y_k):
    # y'_k = (1 − ε_ls) y_k + ε_ls / K
    return (1 - eps_ls) * y_k + eps_ls / K
# the correct token, y_k = 1
round(smooth(1), 6)  # → 0.900003
# every other token, y_k = 0
print(f"{smooth(0):.7f}")  # → 0.0000027

The paper notes that this hurts perplexity, because the model learns to be less sure, but improves accuracy and BLEU. That is a nice example of a training metric and a quality metric disagreeing.

Why it matters today

Very large modern models trained on enormous datasets often use little or no dropout, since they rarely see the same example twice. Label smoothing survives in classification and translation. See the regularization lesson.

6 Results · original

Everyday picture

BLEU scores a translation by how many of its word sequences also appear in a professional human translation, on a 0–100 scale. A gain of 1–2 points is considered a clear improvement. Output was generated with beam search (beam size 4) and averaged the weights from the last few saved checkpoints.

Selected rows of Table 2, reproduced with attribution (Vaswani et al., 2017). BLEU on newstest2014; training cost in FLOPs
ModelBLEU EN→DEBLEU EN→FRTraining cost EN→DE
GNMT + RL24.639.922.3 · 1019
ConvS2S25.1640.469.6 · 1018
Transformer (base)27.338.13.3 · 1018
Transformer (big)28.441.82.3 · 1019

Reading it: each bar is one model's English→German BLEU, with its training cost beside it. The base Transformer beats the previous best single models while costing a fraction of the compute (3.3 × 1018 FLOPs, against 9.6 × 1018 for ConvS2S). The big model adds another point. See the paper for the full table, including ensembles.

What the ablations show (Table 3)

Selected rows of Table 3, reproduced with attribution (Vaswani et al., 2017). EN→DE development set (newstest2013)
VarianthdkPerplexityBLEU
Base model8644.9225.8
(A) one head15125.2924.9
(A) 4 heads41285.0025.5
(A) 16 heads16324.9125.8
(A) 32 heads32165.0125.4
(E) learned positions instead of sinusoids8644.9225.7
Big model16644.3326.4

Rows (A) keep the total width fixed and trade the number of heads against their width. One head is 0.9 BLEU worse than the best setting, and too many narrow heads also hurts: multi-head attention helps, but only up to a point. Row (E) shows that learned positions work about as well as the sine-and-cosine stamps.

Why it matters today

The pattern here, a simpler architecture that is better and cheaper to train, is why the Transformer spread to every modality within three years.

7 Conclusion · original

“In this work, we presented the Transformer, the first sequence transduction model based entirely on attention, replacing the recurrent layers most commonly used in encoder-decoder architectures with multi-headed self-attention.”Vaswani et al. (2017), §7

The authors close by planning to apply attention to images, audio and video, to study restricted (local) attention for long inputs, and to make generation less sequential. All three happened: Vision Transformers treat image patches as tokens, long-context research made windowed and memory-efficient attention standard, and speculative decoding attacks the one-token-at-a-time bottleneck.

What changed since 2017

The core, attention plus feed-forward layers wrapped in residual connections, is unchanged. Nearly every detail around it has been tuned:

Choice in the paperCommon todayWhyLesson
Encoder + decoderDecoder-only for chat models; encoder-only for embeddingsOne stack of masked self-attention generates text; a bidirectional stack makes the best embeddingstransformer
Post-norm: LayerNorm(x + Sublayer(x))Pre-norm, often RMSNormStable training at depthdeep nets
Sinusoidal positionsRoPERelative positions and easier context extensionpositional
ReLU feed-forwardGELU / SwiGLU, sometimes mixture of expertsBetter quality per parametertransformer
8 heads, each with its own K and VGrouped-query attentionA smaller KV cache, so more users per GPUattention
Plain attentionFlashAttention kernelsThe same math with far less memory trafficcompanion
6 layers, 65M–213M parameters30–120 layers, billions of parametersScaling laws: bigger models trained on more data keep improvingcompanion

Glossary

Every term with hover guidance on this page, in one place.