RoFormer: Rotary Position Embedding, annotated
How to read this page
Nothing here assumes you already know the jargon.
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same. Try the θ in the first rotation equation.
- The rotation playground in §3.2.1 is the heart of the page: drag the positions and watch the score stay put.
Each idea climbs the same ladder: an everyday picture, a tiny example you can check by hand, a diagram, the math, then why it matters today. If attention itself is new to you, start with the Attention Is All You Need companion; the code that builds positions from scratch is in the positional information lesson.
Abstract · original
“Specifically, the proposed RoPE encodes the absolute position with a rotation matrix and meanwhile incorporates the explicit relative position dependency in self-attention formulation.”Su et al. (2021), Abstract
Everyday picture
Give every word a clock hand. The first word's hand points at 1 o'clock, the second at 2 o'clock, and so on: that is an absolute position. Now ask a different question: how far apart are two words? You don't need to know either time; you only need the angle between the two hands, and that angle is the same whether the words sit at 1 and 3 o'clock or at 7 and 9. RoPE turns this into mathematics: it rotates each word's query and key by an angle set by its position, so the attention score between two words ends up depending only on how far apart they are.
What the paper claims
- A position scheme that stores absolute positions but produces relative behaviour: attention scores depend on the gap between words, not on where the pair sits.
- Three useful properties: it works for any sequence length (no learned table with a fixed number of rows), word pairs that are far apart get a naturally weaker link, and it is compatible with linear attention.
- The resulting model, RoFormer, trains faster than its baselines and does better on long documents.
Why it matters today
RoPE became the default position scheme for large language models: the Llama family, Mistral and many others rotate queries and keys exactly as described here. Understanding this paper explains a line of code that runs in nearly every open model, and why "extending the context window" of a model is mostly a question about rotation speeds.
1 Introduction · original
“The sequential order of words is of great value to natural language understanding.”Su et al. (2021), §1
Everyday picture
“Dog bites man” and “man bites dog” use the same three words; only the order changes the news. A recurrent network gets order for free, because it reads one word at a time. A Transformer reads all words at once, like tipping a bag of Scrabble tiles onto the table: without extra help, it cannot tell which tile came first. That property has a name, permutation equivariance: shuffle the input words and the outputs simply shuffle the same way.
Tiny example
Take three word vectors and run plain attention (no positions). Swap the first and last word. Every score is a dot product between two words' vectors, and a dot product does not care where the words sit, so the score grid is the same numbers in shuffled rows and columns. The model literally cannot tell the two sentences apart. Position information has to be injected somehow, and this paper is about the best way to do it.
Why it matters
Every Transformer needs some position scheme. The choice affects quality, how well the model copes with texts longer than it saw in training, and how cheaply the context window can later be stretched. The positional information lesson demonstrates the “dog bites man” problem in code.
2 Background · original
2.1 Queries and keys that know their position · original
Everyday picture
Before a word joins the attention step, it prepares three cards: a question (query), a label (key) and a message (value). The only question in this whole debate is: when and how does the word's seat number get written onto those cards?
In words: “each word's query, key and value are made from two ingredients: the word itself and the position it sits at.”
With the numbers: in “the cat sat”, the query for “sat” is fq(vector for “sat”, 3), and the key for “the” is fk(vector for “the”, 1). Different position schemes are simply different recipes for f.
In Python:
sentence = ["the", "cat", "sat"]
# each word with its seat m
for m, x_m in enumerate(sentence, start=1):
print(f"q_{m} = f_q({x_m!r}, {m})") # → q_1 = f_q('the', 1) q_2 = f_q('cat', 2) q_3 = f_q('sat', 3)
2.2 Absolute position embedding · original
Everyday picture
The classic recipe staples a seat-number badge onto each word before anything else happens: add a position vector pi to the word vector xi, then make the query, key and value from the sum. The original Transformer's badges were fixed sine and cosine waves; BERT and GPT-2 instead learned one badge per seat, up to a maximum number of seats.
In words: “add the seat's badge to the word, then project the sum into a query, key or value.”
With the numbers: with 2-number toy vectors, word x = (1, 0) at seat 3 with badge p3 = (0.1, 0.9) becomes (1.1, 0.9) before projection. The same word at seat 4 would carry a different badge, so it enters attention as a slightly different vector.
In Python:
# the word's vector
x = [1, 0]
# the badge for seat 3
p_3 = [0.1, 0.9]
# x_i + p_i, before W_t projects it
[x_k + p_k for x_k, p_k in zip(x, p_3)] # → [1.1, 0.9]
The catch
Badges are added before the words meet, so the score between two words mixes up four things: word with word, word with badge, badge with word, and badge with badge. The relative distance is only implicitly buried in there. Learned badges also stop at the last seat the model was trained with: a model with 512 learned badges has nothing to say about seat 600.
2.3 Relative position embedding · original
Everyday picture
Instead of seat numbers, tell each pair of words how far apart they are: “three seats to your left”. Shaw et al. (2018) did this by adding a learned vector for each distance (clipped, so everything beyond some distance shares one vector); Transformer-XL and T5 found other ways to add a distance-dependent term to the score.
The catch
These schemes bolt an extra term onto the attention computation, which means extra memory, extra code in the attention kernel, and trouble combining with faster attention variants that never compute the full grid of scores. The paper surveys several and asks: can we get purely relative behaviour without changing attention at all?
Why it matters
This is the design space RoPE lands in: the simplicity of absolute schemes (just transform q and k before attention) with the behaviour of relative ones (scores depend only on distance).
3 Proposed approach · original
3.1 The goal, written as an equation · original
“In other words, we hope that the inner product encodes position information only in the relative form.”Su et al. (2021), §3.1
Everyday picture
Two people shaking hands should feel the same handshake whether they meet on the first floor or the fortieth. The strength of the handshake (the attention score) may depend on who they are and how far apart they stood, never on which floor the meeting happened on.
In words: “find recipes for the query and the key such that their dot product depends on the two words and on the gap between their positions, and on nothing else.”
With the numbers: the score between a query at seat 2 and a key at seat 1 must equal the score for the same two words at seats 7 and 6, because m − n = 1 in both cases.
In Python:
m, n = 2, 1
m - n # → 1
m, n = 7, 6
# the same gap, so g must give the same score
m - n # → 1
Why it matters
Writing the wish as an equation turns a vague goal into a puzzle with a clean answer. The rest of §3 solves it.
3.2.1 The 2D case: a rotation · original
Everyday picture
Draw the query and key as arrows on a sheet of paper. Rotate the query arrow by an angle proportional to its position, and the key arrow by an angle proportional to its own position. The dot product of two arrows depends only on their lengths and the angle between them. Rotating both arrows by the same extra amount leaves that angle unchanged. So if both words move three seats to the right, both arrows turn by the same extra angle, and the score cannot change.
Tiny example
Take the simplest arrows, q = (1, 0) and k = (1, 0), and turn one radian per position (θ = 1). A query at position 2 turns by 2 radians and a key at position 1 by 1 radian; the angle between them is 1 radian, so the score is cos(1) = 0.540. Move the pair to positions 7 and 6: the arrows turn by 7 and 6 radians, the gap is still 1 radian, and the score is still 0.540. Move the key back to position 1 while the query stays at 3: now the gap is 2 radians and the score drops to cos(2) = −0.416.
Reading it: the faint arrows are the query and key before rotation; the bold arrows are after. The blue arrow is the query at position m, turned by m·θ; the orange arrow is the key at position n, turned by n·θ. The shaded wedge is the angle between them, and the readout shows the resulting attention score, the dot product. Press “Shift both positions by +1” a few times: both arrows swing round, but the wedge keeps its size and the score does not move. Now drag only m: the wedge changes, and so does the score. That is the whole paper in one picture.
Hover or tap a block. Start at the bottom with the word vector.
Reading it: read upwards. The word vector is projected into a query exactly as in any Transformer. RoPE then chops the query into neighbouring pairs of numbers and treats each pair as an arrow on a flat plane. Each arrow is turned by an angle: the word's position m times that pair's own speed θi. The turned pairs are glued back together into the position-encoded query. Nothing is added and nothing is learned in the rotation step: position enters purely as an angle.
The math
In words: “make the query (or key) as usual with a learned matrix, then turn the resulting 2D arrow by position times θ.”
With the numbers: with θ = 1 and m = 2, the matrix is [[cos 2, −sin 2], [sin 2, cos 2]] = [[−0.416, −0.909], [0.909, −0.416]], which turns q = (1, 0) into (−0.416, 0.909): the same arrow, pointing 2 radians further round.
In Python:
import math
theta, m = 1, 2
R = [[math.cos(m * theta), -math.sin(m * theta)],
[math.sin(m * theta), math.cos(m * theta)]]
[[round(v, 3) for v in row] for row in R] # → [[-0.416, -0.909], [0.909, -0.416]]
q = [1, 0]
# R q
[round(sum(R[i][j] * q[j] for j in range(2)), 3) for i in range(2)] # → [-0.416, 0.909]
Why does this solve §3.1? A rotation by angle a, followed by the transpose of a rotation by angle b, is a single rotation by b − a. So the score becomes
In words: “rotating the query by its position and the key by its position is the same as leaving the query alone and rotating the key by the difference of the positions.”
With the numbers: for q = (2, 1) and k = (1, 1) with θ = 1, positions (3, 1) and positions (8, 6) both give a score of −0.339, because both pairs have n − m = −2.
In Python:
import math
# turn a 2D arrow by angle
def R(angle, v):
c, s = math.cos(angle), math.sin(angle)
return [c * v[0] - s * v[1], s * v[0] + c * v[1]]
def dot(a, b):
return sum(a_k * b_k for a_k, b_k in zip(a, b))
q, k, theta = [2, 1], [1, 1], 1
# (R_m q)ᵀ (R_n k), m = 3, n = 1
round(dot(R(3 * theta, q), R(1 * theta, k)), 3) # → -0.339
# m = 8, n = 6
round(dot(R(8 * theta, q), R(6 * theta, k)), 3) # → -0.339
# qᵀ R_(n−m) k
round(dot(q, R((1 - 3) * theta, k)), 3) # → -0.339
Why it matters today
This two-line argument is the reason RoPE needs no extra parameters, no change to the attention kernel, and no table of learned distances. It is also why it composes with KV caching: a key is rotated once, when it is written, and never touched again. See the positional information lesson, which checks this property numerically.
3.2.2 General form: a wall of clocks · original
Everyday picture
One rotating arrow has a problem: after a full turn it repeats, so positions 0 and 6.28 look identical at θ = 1. The fix is the same trick the original Transformer used for its position badges, and the same trick a clock uses: have many hands turning at very different speeds. The seconds hand distinguishes nearby moments, the hours hand distinguishes distant ones, and together they are never ambiguous. RoPE splits a query of d numbers into d/2 pairs and gives each pair its own speed.
Tiny example
With a toy width of d = 8 there are 4 pairs, and the paper's speed rule gives θ = 1, 0.1, 0.01 and 0.001 radians per position. At position 10, the four hands have turned by 10, 1, 0.1 and 0.01 radians: the fast hand has already gone round more than once, while the slowest has barely moved.
Reading it: each dial is one pair of numbers in the query (d = 8, so four pairs). The hand shows how far that pair has been turned at the position chosen on the slider. Drag the slider slowly: the left dial spins quickly and repeats every 6.3 positions, so it can only tell nearby positions apart; the right dial barely moves until you reach thousands of positions, so it tracks coarse distance. Reading all four together pins the position down exactly, the same way hours, minutes and seconds together tell the time.
In words: “rotate the first pair by m·θ1, the second pair by m·θ2, and so on, where the speeds shrink geometrically from 1 radian per position down to about 1/10,000.”
With the numbers: for d = 8: θ1 = 100000 = 1, θ2 = 10000−2/8 = 0.1, θ3 = 10000−4/8 = 0.01, θ4 = 10000−6/8 = 0.001.
In Python:
d = 8
# θ_1 … θ_(d/2)
[round(10000 ** (-2 * (i - 1) / d), 3) for i in range(1, d // 2 + 1)] # → [1.0, 0.1, 0.01, 0.001]
The full score then reads : the two position rotations merge into one rotation by the gap, in every pair at once. The paper also notes that the matrix is almost entirely zeros, so nobody ever builds it (§3.4.2).
Why it matters today
The base 10,000 sets the slowest clock, and therefore the longest distance the model can tell apart cleanly. Context-extension tricks such as position interpolation, NTK-aware scaling and YaRN all work by adjusting these speeds, which is why RoPE models can be stretched to longer contexts after training.
3.3 Properties of RoPE · original
Everyday picture
Three promises come with the design. No seat limit: a rotation angle can be computed for position 50,000 as easily as for position 5, so there is no learned table to run off the end of. Distance fades: words far apart naturally interact more weakly (the plot in §3.4.3 shows why). Plays well with shortcuts: because position lives inside q and k rather than in an extra score term, it survives inside linear-attention variants that never build the full score grid.
The paper demonstrates the last point with the Performer, a linear-attention model, in §4.4.
Why it matters
“No seat limit” is subtle: RoPE can compute any angle, but a model trained only on short texts has never seen the large angles the slow clocks reach at long range. That gap is exactly what later context-extension methods patch up.
3.4.2 The fast way to compute it · original
Everyday picture
A d × d rotation matrix for d = 4096 would have 16.7 million entries, almost all zero. Nobody multiplies by all those zeros. Rotating an arrow (a, b) by angle φ only needs four multiplications: (a cos φ − b sin φ, b cos φ + a sin φ). Do that for every pair at once and you have RoPE.
In words: “multiply every number by the cosine of its pair's angle, then add the pair-swapped, sign-flipped copy multiplied by the sine.”
With the numbers: x = (1, 2, 3, 4) at position m = 1 with θ = (1, 0.01). Cosines (0.540, 0.540, 1.000, 1.000), sines (0.841, 0.841, 0.010, 0.010), rot(x) = (−2, 1, −4, 3). Result: (1·0.540 − 2·0.841, 2·0.540 + 1·0.841, 3·1.000 − 4·0.010, 4·1.000 + 3·0.010) = (−1.143, 1.922, 2.960, 4.030). The fast pair has visibly turned; the slow pair has barely moved.
In Python:
import math
x, m, theta = [1, 2, 3, 4], 1, [1, 0.01]
# both numbers in a pair share mθ
angle = [m * t for t in theta for _ in (0, 1)]
# swap each pair, flip a sign
rot_x = [-x[1], x[0], -x[3], x[2]]
[round(x_k * math.cos(a) + r_k * math.sin(a), 3) for x_k, r_k, a in zip(x, rot_x, angle)] # → [-1.143, 1.922, 2.96, 4.03]
Why it matters today
This is, almost line for line, the “rotate half” function found in open-source model code. The cost is a few element-wise multiplications per query and key, which is negligible next to the attention itself.
3.4.3 Long-term decay · original
Everyday picture
Picture many clock hands, all starting together at 12 o'clock. Nudge the time forward a little and they still point roughly the same way: add them up as arrows and you get a long arrow. Nudge it forward a lot and the hands point in every direction; added up, they mostly cancel. The score between two words is built from exactly such a sum of rotated contributions, one per pair, so for words that are far apart the contributions tend to cancel and the possible score shrinks.
In words: “add up the first j clock hands as arrows and measure the length of the result; average those lengths over all j. The paper proves the attention score's size is capped in proportion to this average.”
With the numbers: for d = 128 at distance 0 every hand points the same way, so |Sj| = j and the average of 1, 2, …, 64 is 32.5. At distance 10 it has fallen to about 18.0, and at distance 100 to about 10.2.
In Python:
import cmath
d = 128
# θ_0 … θ_(d/2 − 1)
theta = [10000 ** (-2 * i / d) for i in range(d // 2)]
# gap = m − n
def bound(gap):
S = [sum(cmath.exp(1j * gap * theta[i]) for i in range(j)) for j in range(1, d // 2 + 1)]
return sum(abs(S_j) for S_j in S) / (d / 2)
round(bound(0), 1), round(bound(10), 1), round(bound(100), 1) # → (32.5, 18.0, 10.2)
Hover or tap the curve to read the bound at any distance.
Reading it: the x-axis is the distance between two words (m − n) and the y-axis is the paper's upper bound on how large their attention score can be, relative to the size of the vectors. It falls fast over the first few dozen positions and keeps drifting down with wiggles. The model can still attend strongly to a distant word if the learned vectors line up, but the default leans towards nearby words, which matches how language usually works.
Why it matters today
The decay is a gentle prior, not a hard rule, and later work showed the picture is subtler in trained models. It also explains a practical failure: past the lengths seen in training, the slow clocks enter angles the model never learned, and quality collapses unless the speeds are rescaled.
4 Experiments · original
Everyday picture
A new part for an engine is tested by swapping it into several existing engines and running them on the usual tracks. The paper swaps RoPE into a translation Transformer, into BERT, and into the Performer, and tries long Chinese legal documents.
| Test | Baseline | Baseline | RoFormer |
|---|---|---|---|
| WMT 2014 English→German translation (BLEU, §4.1) | Transformer-base | 27.3 | 27.5 |
| GLUE: MRPC (F1, §4.3) | BERT | 88.9 | 89.5 |
| GLUE: STS-B (Spearman correlation) | BERT | 85.8 | 87.0 |
| GLUE: QQP (F1) | BERT | 71.2 | 86.4 |
| GLUE: SST-2 (accuracy) | BERT | 93.5 | 90.7 |
| GLUE: QNLI (accuracy) | BERT | 90.5 | 88.0 |
| GLUE: MNLI matched / mismatched (accuracy) | BERT | 84.6 / 83.4 | 80.2 / 79.8 |
| CAIL2019-SCM long legal texts, test accuracy (§4.5) | WoBERT, 512 tokens | 68.10% | 69.79% (1024 tokens) |
Read the table honestly. On translation the gain is small (0.2 BLEU). On GLUE, RoFormer wins three tasks and loses three, which the paper summarises as a significant win on three of six. The clearest result is the last row: with 1,024-token inputs RoFormer beats the 512-token models by 1.5 to 2 points on long documents, the setting where a position scheme matters most. Pre-training plots (§4.2, §4.4) show RoFormer's loss falling faster than BERT's and the Performer's, but give no final numbers.
Why it matters today
RoPE did not win on headline benchmark gains. It won on engineering: no parameters, no kernel changes, relative behaviour, and a clean handle for stretching context. Those properties mattered more and more as models and context windows grew.
4.5.5 Limitations and 5 Conclusions · original
“Our theoretical analysis indicates that relative position can be naturally formulated using vector production in self-attention, with absolution position information being encoded through a rotation matrix.”Su et al. (2021), §5
The authors are candid about two open questions: they cannot fully explain why RoFormer converges faster than baselines, and although they prove the decay property, it is shared by other schemes, so it does not by itself explain the better long-text results.
What happened next
| Development | What it does to the clocks | Where to learn more |
|---|---|---|
| Adoption in large language models | RoPE, usually applied to every layer's queries and keys, became the default in open LLMs such as the Llama family | positional lesson |
| Position interpolation | Slows every clock so long inputs map into the range of angles seen in training | positional lesson |
| NTK-aware scaling and YaRN | Slow the slow clocks more than the fast ones, keeping local word order sharp | positional lesson |
| ALiBi | An alternative with no clocks at all: subtract a penalty that grows with distance | positional lesson |
Glossary
Every term with hover guidance on this page, in one place.