Neural Machine Translation of Rare Words with Subword Units, annotated
How to read this page
- Any dotted word explains itself on hover, focus or tap; so does every symbol in every equation.
- The BPE machine in §3.2 is live: step through the merges on the paper's own example, or type your own text and watch a vocabulary grow from letters.
Each idea climbs the ladder: everyday picture, tiny example, diagram, math, why it matters today. The tokenization lesson builds BPE from scratch in Python and shows its quirks.
Abstract · original
“Neural machine translation (NMT) models typically operate with a fixed vocabulary, but translation is an open-vocabulary problem.”Sennrich, Haddow and Birch (2015), Abstract
Everyday picture
A neural translator in 2015 had a fixed word list, around 30,000 to 50,000 entries. Any word not on the list became a blank, “UNK” (unknown). But language is endless: new names, new compounds, rare inflections. The paper's fix: instead of whole words, use a fixed kit of word pieces, like a Lego set. Common words are single bricks; rare words are built from several. Then no word is ever unknown.
What the paper claims
- Rare and unknown words can be translated by breaking them into subword units, with no separate dictionary for unknown words.
- Byte pair encoding, an old compression trick, makes a good way to choose the pieces.
- Better translation of rare words, and gains of up to 1.1 and 1.3 BLEU over a dictionary-based baseline for English→German and English→Russian.
Why it matters today
GPT, Claude, Llama and most other language models split text with a descendant of this algorithm (usually byte-level BPE). Why models are priced per token, why they struggle to count letters, and why some languages cost more than others all trace back to here.
1 Introduction · original
Everyday picture
German glues words together: Abwasserbehandlungsanlage is “sewage water treatment plant” as a single word. A word-level model has either seen that exact word or it hasn't. A human translator splits it into parts they know, Abwasser | behandlungs | anlage, and translates the parts. The paper wants the network to do the same.
Tiny example
Before this paper, the standard patch was a back-off dictionary: when the model outputs UNK, look up the aligned source word in a bilingual dictionary, or copy it across. That works for “Obama” into German, but not for “Obama” into Russian, which needs a different alphabet (Обама), and not for compounds that have no one-to-one counterpart.
Two contributions
- Open-vocabulary translation works by encoding rare words as sequences of subword units, and it is simpler and better than large vocabularies plus back-off dictionaries.
- Byte pair encoding, a compression algorithm from 1994, adapted to choose those units.
Why it matters today
“Never have an unknown token” became a hard requirement for language models; modern tokenizers go further, down to raw bytes, so even emoji and binary junk have a spelling.
2 Neural machine translation · original
Everyday picture
The translator used is the attention-based encoder-decoder of 2015: a recurrent encoder reads the source sentence in both directions, and a recurrent decoder writes the translation one piece at a time, using attention to look back at the most relevant source pieces at each step. The paper stresses that its method is not tied to this architecture: it only changes what the pieces are.
Tiny example
With subwords, attention can focus on part of a word: when writing Gesundheits, it can attend to the English piece “health”, and when writing forschungs, to “research”. A word-level model can only attend to whole words.
Setup, for reference
Hidden layers of 1,000 units, embeddings of 620, the Adadelta optimizer with mini-batches of 80, about 7 days of training per model, and beam search with a beam of 12.
Why it matters today
Two years later the Transformer replaced the recurrent parts, and it adopted subword units from the start: the Transformer paper used byte pair encoding for English→German.
3 Subword translation · original
Everyday picture
A good human translator can handle a word they have never seen if they know its parts. The paper names three kinds of word where that works: names (copy or transliterate them), cognates and loanwords (translate with regular letter changes), and morphologically complex words (translate each part).
| Kind | English | Other languages |
|---|---|---|
| Name | Barack Obama | Барак Обама (Russian), バラク・オバマ (Japanese) |
| Cognate / loanword | claustrophobia | Klaustrophobie (German), Клаустрофобия (Russian) |
| Compound | solar system | Sonnen + system (German), Nap + rendszer (Hungarian) |
Tiny example: what rare words really are
The authors examined 100 rare German words from their training data (words outside the 50,000 most frequent). 56 were compounds, 21 names, 6 loanwords, 5 transparent affixations (like süßlich = süß + lich, “sweetish”), 1 number and 1 computer-language identifier. So the great majority could, in principle, be translated piece by piece.
Reading it: each bar counts how many of the 100 sampled rare German words fell into each category. Compounds dominate: over half of all rare words were glued-together combinations of words the model may well know. Together with names, that is 77 of the 100, all candidates for piece-by-piece translation. Numbers from §3 of the paper.
Why it matters today
The same argument applies with more force to code (identifiers like parseHttpResponseHeader) and to morphologically rich languages; subword tokenizers let one model handle all of them.
3.1 Related work · original
Everyday picture
There is a trade-off between two costs. A big vocabulary of whole words makes sentences short (few tokens) but the model huge, and rare words starve for training data. A tiny vocabulary of characters makes the model small but sentences very long, so every step of reasoning must span many more positions. The authors want a middle point: a compact vocabulary that still keeps text short.
Tiny example
“unbelievable” as whole words: 1 token, from a vocabulary of millions. As characters: 12 tokens, from a vocabulary of about 100. As subwords: perhaps “un” + “believ” + “able”, 3 tokens, from a vocabulary of tens of thousands.
Why it matters today
Choosing the vocabulary size is still exactly this trade-off, now weighed against GPU memory for the embedding table and the cost of long sequences in attention, which grows with the square of length.
3.2 Byte pair encoding (BPE) · original
“We iteratively count all symbol pairs and replace each occurrence of the most frequent pair (‘A’, ‘B’) with a new symbol ‘AB’.”Sennrich, Haddow and Birch (2015), §3.2
Everyday picture
Start with a box of single letters. Look through all your text and find the two pieces that sit side by side most often, perhaps “e” followed by “s”. Glue them into one new piece, “es”, and add it to the box. Repeat: maybe “es” + “t” is now the most common pair, giving “est”. After a few thousand repeats, common words are single pieces, frequent chunks (“ing”, “tion”) are pieces, and rare words are spelled from several. The number of repeats is the only setting.
Tiny example: the paper's own toy dictionary
Four words with their counts: low (5), lower (2), newest (6), widest (3). Each word is split into characters plus an end-of-word marker </w>, which lets you glue words back together after translation. Count the pair (e, s): it appears once in “newest” (6 occurrences) and once in “widest” (3 occurrences), so 6 + 3 = 9. Nothing beats 9, so the first merge is e + s → es. The next is es + t → est (also 9), then est + </w> (9), then l + o (5 + 2 = 7), then lo + w (7). Step through it below.
Most frequent adjacent pairs right now
Merges learned so far, in order
Reading it: the top row shows every word in the corpus, split into its current pieces, with its count; newly glued pieces are highlighted. The left table counts every pair of neighbouring pieces, weighted by word counts; the highlighted row is the winner, which the next merge will glue. The right table is the growing list of learned merges, the tokenizer itself. Press Next merge repeatedly on the paper's Algorithm 1 corpus and you get the same ten merges as the paper's printed code: es, est, est</w>, lo, low, ne, new, newest</w>, low</w>, wi. Then try the segmenter: after those merges, “lowest”, a word the corpus never contained, comes out as low + est</w>, and “lower” as low + e + r + </w>: no word is ever unknown, because in the worst case it falls back to letters. On the Figure 1 corpus every pair ties at first, so the order of merges is decided by tie-breaking; the paper's figure shows one valid order and this demo may pick another.
| Merge | Result |
|---|---|
| r + · | r· |
| l + o | lo |
| lo + w | low |
| e + r· | er· |
(The paper's figure writes the end-of-word marker as “·” and its code as “</w>”; they are the same thing.) With those four merges, the unseen word “lower” is segmented as low er·.
The math
In words: “the next merge is the neighbouring pair that occurs most often, counting each word as many times as it appears; the final vocabulary is the starting characters plus one new symbol per merge.”
With the numbers: for (e, s): f(newest) · 1 + f(widest) · 1 = 6 + 3 = 9, the maximum, so (A, B) = (e, s). The paper's German BPE used M = 59,500 merges; with the few hundred starting characters, that gives the roughly 60,000-symbol vocabulary of Table 1.
In Python:
# f(w)
f = {"l o w </w>": 5, "l o w e r </w>": 2, "n e w e s t </w>": 6, "w i d e s t </w>": 3}
counts = {}
# Σ over words w ...
for w, f_w in f.items():
symbols = w.split()
# ... each neighbouring pair, n_ab(w) times ...
for a, b in zip(symbols, symbols[1:]):
# ... weighted by f(w)
counts[(a, b)] = counts.get((a, b), 0) + f_w
# 6 + 3
counts[("e", "s")] # → 9
# argmax: (s, t) also scores 9; the first found wins the tie
max(counts, key=counts.get) # → ('e', 's')
# "a few hundred" characters (500 here) plus the paper's merges
V_chars, M = 500, 59_500
# |V| = |V_chars| + M
V_chars + M # → 60000
Two practical details from the paper: pairs are never counted across word boundaries, so the algorithm can run on a word-frequency list instead of raw text; and applying the merges to new text simply replays them in the order they were learned. The paper also compares learning separate merges for each language with joint BPE, one set learned on both languages together (with Russian transliterated into Latin letters for learning), which keeps names segmented consistently on both sides.
Why it matters today
This loop, count pairs, merge the top one, repeat, is essentially how the tokenizers of GPT-style models are trained today, just on bytes instead of characters and over far more text. See the tokenization lesson, which reproduces the “low lower lowest” merges exactly.
4 Evaluation · original
Everyday picture
Two questions: do subwords translate rare words better, and which way of splitting is best for vocabulary size, text length and quality? The test bed is WMT 2015: English→German (4.2 million sentence pairs, about 100 million tokens) and English→Russian (2.6 million pairs, about 50 million tokens).
How quality is measured
BLEU counts overlapping word sequences with a reference translation; chrF3 does the same with character sequences, weighting recall more. Because the paper's claim is about rare words, it also reports F1 on single words, separately for all words, rare words and words never seen in training (OOV).
In words: “F1 is the harmonic mean of precision and recall, high only when both are; F-beta tilts that balance: β = 3 (chrF3) treats recall as three times as important as precision, which is why β² = 9 appears in the formula.”
With the numbers: the joint-BPE system's English→German OOV words had precision 38.6% and recall 29.8%, so F1 = 2 × 0.386 × 0.298 / (0.386 + 0.298) = 0.336, the 33.6% in Table 2. With β = 3 the same precision and recall give F3 = 10 × 0.386 × 0.298 / (9 × 0.386 + 0.298) = 0.305, pulled towards the lower of the two, recall.
In Python:
# precision and recall
P, R = 0.386, 0.298
F1 = 2 * P * R / (P + R)
round(F1, 3) # → 0.336
# chrF3's beta: recall counts three times as much
beta = 3
F_beta = (1 + beta ** 2) * P * R / (beta ** 2 * P + R)
round(F_beta, 3) # → 0.305
4.1 Subword statistics · original
Everyday picture
Every way of splitting text trades vocabulary size (how many different pieces) against text length (how many pieces in total). The ideal sits in a corner: few types, short text, and zero unknown pieces on new text.
Hover or tap a point to see the segmentation it represents.
Reading it: each point is one way of splitting the same 100-million-word German corpus. Left is shorter text, down is a smaller vocabulary; points coloured red still leave unknown tokens on the test set. Whole words (top left) keep text short but need 1.75 million types and still hit 1,079 unknowns. Characters (bottom right) need only 3,000 types and have no unknowns, but make the text 5.5× longer. Classic linguistic splitters (compound splitting, Morfessor, hyphenation) barely shrink the vocabulary and all leave unknowns. BPE sits near the ideal corner: 63,000 types, text only 12% longer than words, and zero unknown tokens.
| Segmentation | # tokens | # types | # UNK |
|---|---|---|---|
| none (whole words) | 100 m | 1,750,000 | 1,079 |
| characters | 550 m | 3,000 | 0 |
| character bigrams | 306 m | 20,000 | 34 |
| character trigrams | 214 m | 120,000 | 59 |
| compound splitting | 102 m | 1,100,000 | 643 |
| Morfessor | 109 m | 544,000 | 237 |
| hyphenation | 186 m | 404,000 | 230 |
| BPE | 112 m | 63,000 | 0 |
| BPE (joint) | 111 m | 82,000 | 32 |
| character bigrams, 50,000-word shortlist | 129 m | 69,000 | 34 |
Why it matters today
The same trade-off is why a token is about ¾ of an English word today, and why text in languages the tokenizer saw less of splits into more tokens and so costs more.
4.2 Translation experiments · original
Everyday picture
The contest: word-level models with a huge vocabulary, with and without a back-off dictionary (WUnk, WDict), against subword models with no dictionary at all (character bigrams C2-50k; BPE-60k; joint BPE-J90k).
Tiny example
For English→German rare words (not in the top 50,000), unigram F1 rises from 36.8% (word model with dictionary) to 41.8% (joint BPE). For English→Russian, where names must be transliterated, OOV-word F1 rises from 6.6% to 18.3%: the dictionary can copy “Mirzayeva” but not turn it into Cyrillic.
| System | EN→DE BLEU | EN→DE rare F1 | EN→DE OOV F1 | EN→RU BLEU | EN→RU rare F1 | EN→RU OOV F1 |
|---|---|---|---|---|---|---|
| WUnk (words, UNK) | 22.8 | 20.4 | 0.0 | 22.4 | 25.2 | 0.0 |
| WDict (words + dictionary) | 24.2 | 36.8 | 36.8 | 22.8 | 26.5 | 6.6 |
| C2-50k (char bigrams) | 25.3 | 40.5 | 30.9 | 24.1 | 27.8 | 17.4 |
| BPE-60k | 24.5 | 40.9 | 29.3 | 23.6 | 29.7 | 15.6 |
| BPE-J90k (joint) | 24.7 | 41.8 | 33.6 | 24.1 | 29.7 | 18.3 |
Reading it: each group compares the word-level system with a dictionary (grey) against joint BPE (blue) on single-word F1 for one kind of word. For all words the two are nearly tied, because frequent words are single tokens either way. The gap opens for rare words, and for English→Russian unseen words it nearly triples, because transliterating names needs subword pieces. English→German OOV is the one place the dictionary wins: most unseen German words are names that can simply be copied. Numbers from Tables 2 and 3 of the paper.
Rare and unseen words make up only 9 to 11% of the test sets, so the headline BLEU gains are modest: the subword ensembles beat the dictionary baseline by 0.3 to 1.3 BLEU. The authors argue that BLEU underestimates the benefit, because rare words (names, key terms) tend to carry a sentence's central information.
Why it matters today
“Better on the long tail, same on the head” is the typical signature of a tokenization improvement, and why average metrics can hide it.
5 Analysis · original
Everyday picture
Two checks. First, sort target words by how often they appeared in training and plot accuracy for each band: word-level models collapse on rare words, subword models degrade gracefully. Second, look at actual translations.
| System | “health research institutes” → |
|---|---|
| reference | Gesundheitsforschungsinstitute |
| WDict | Forschungsinstitute (drops “health”) |
| C2-50k | Fo|rs|ch|un|gs|in|st|it|ut|io|ne|n |
| BPE-60k | Gesundheits|forsch|ungsinstitu|ten |
| BPE-J90k | Gesundheits|forsch|ungsin|stitute |
Tiny example
The subword models build a correct German compound from pieces, even though the pieces do not follow linguistic boundaries: “forsch | ungsin | stitute” is not how a linguist would split it, yet the translation comes out right. For Russian, joint BPE segments the name Mirzayeva consistently on both sides (Mir | za | yeva → Мир | за | ева), while separate BPE vocabularies occasionally learned spurious letter mappings from inconsistent splits.
Why it matters today
Tokens are not morphemes, and models cope anyway. That is both why BPE works in practice and why models behave oddly on tasks that depend on individual letters: they never see the letters directly.
6 Conclusion · original
“The main contribution of this paper is that we show that neural machine translation systems are capable of open-vocabulary translation by representing rare and unseen words as a sequence of subword units.”Sennrich, Haddow and Birch (2015), §6
The authors note that their choice of vocabulary size was “somewhat arbitrary”, and that shrinking the vocabulary and representing more words as subwords can actually improve performance. The code was released as subword-nmt.
What changed since 2015
| In the paper | Common today | Why | Where to learn more |
|---|---|---|---|
| BPE over characters | Byte-level BPE over the raw bytes of UTF-8 text | Any string, including emoji and code, can be encoded with no unknown symbol at all | tokenization lesson |
| Text pre-split into words; merges never cross word boundaries | A regular-expression pre-tokenizer (spaces, digits, punctuation), or none at all with SentencePiece | Language-independent, and handles languages written without spaces | SentencePiece paper |
| Greedy merges by frequency | Also: the Unigram model, which prunes a large vocabulary by likelihood (Kudo, 2018) | Gives several possible segmentations, which can be sampled as a regularizer | subword regularization paper |
| Vocabularies of 60,000 to 90,000 | 32,000 to about 200,000 tokens, shared across many languages | Bigger vocabularies shorten text, which saves attention compute and API cost | tokenization lesson |
Glossary
Every term with hover guidance on this page, in one place.