An annotated companion · AI Primer

Retrieval-Augmented Generation, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The equations are reproduced because mathematics is not copyrightable. The paper is distributed under arXiv's standard licence, so its tables are not reproduced; instead a few key numbers appear in this page's own tables, each attributed. Figure 1 is redrawn from scratch. Read the original alongside: every section links to it.

How to read this page

Nothing here assumes you already know the jargon.

  • Any dotted word explains itself when you hover, tab to, or tap it.
  • Every symbol inside an equation does the same, and each equation has a table decoding its symbols.
  • The diagram and the demos are live. The step-through in §4 walks one question through the whole system.

Each idea climbs the same ladder: an everyday picture, a tiny example you can check by hand, a diagram, then the math, then why it matters today. The code version lives in the RAG lesson.

Abstract

“We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG).”Lewis et al. (2020), Abstract. Read the original

Everyday picture

A closed-book exam tests what you memorised. An open-book exam lets you look things up first. A plain language model sits every exam closed-book: everything it knows is baked into its parameters. This paper hands the model a library (all of Wikipedia, cut into short passages) and a librarian who fetches the most relevant pages before it writes each answer. That combination is retrieval-augmented generation.

What the paper claims

  • A single recipe that joins a pre-trained retriever and a pre-trained text generator, then fine-tunes them together on any input-to-output task.
  • New best results on three open-domain question-answering benchmarks, beating both models that answer from memory and specialised “find the answer span” systems.
  • More factual, more specific and more varied text than the same generator without retrieval.
  • Knowledge you can update by swapping the library, without retraining.

Why it matters today

Almost every assistant that answers from a company's own documents is a descendant of this idea: fetch first, then write with the fetched text in view, and cite it. The details have changed (see What changed since 2020), but the split between what the model knows and what it can look up comes straight from here.

1 Introduction · original

“They cannot easily expand or revise their memory, can’t straightforwardly provide insight into their predictions, and may produce ‘hallucinations’.”Lewis et al. (2020), §1, on models that store knowledge only in their weights

Everyday picture

Think of two kinds of memory. Parametric memory is what you know by heart: fast, fluent, but impossible to inspect and awkward to update. Non-parametric memory is a filing cabinet: slow to search, but you can open a drawer, show someone the page, and replace an outdated sheet. The paper argues a good system needs both, and joins them so each covers the other's weakness.

Tiny example

Ask “Who is the President of Peru?” A model trained on 2018 text answers with whoever held office in 2018, and nothing short of retraining changes that. A retrieval model answers from whatever passage the filing cabinet returns, so replacing the cabinet's pages with 2016 pages changes the answer to the 2016 office holder. §4.5 of the paper runs exactly this experiment.

The three problems it targets

Problem with memory-only modelsHow retrieval helps
Knowledge is frozen at training timeEdit or swap the document index
No way to see why it said somethingThe retrieved passages are readable evidence
Hallucination: confident, unsupported claimsGeneration is conditioned on real text, so it drifts less

Why it matters today

These three problems are still the reasons teams reach for RAG: freshness, grounding with citations, and fewer invented facts.

2 Methods · original

Everyday picture

A research assistant gets your question, turns it into a search, pulls the five most promising index cards from a card catalogue of 21 million cards, and hands them to a writer. The writer drafts an answer while reading the cards. Crucially, the writer does not have to trust any single card: it weighs its answer across all of them, leaning on the ones the assistant was most confident about.

Hover or tap any block in the redrawn figure. The numbers use the paper's setup: K = 5 retrieved passages from an index of 21 million 100-word Wikipedia chunks.

Retriever p_η(z | x) · non-parametric memory Generator p_θ · parametric memory x QueryEncoder MIPStop-K DocumentIndex d(z) Generator(BART)once per z Σ overz → y z₁ … z_K x is passed to the generator too

Hover or tap a block. Start at x on the left and follow the arrows.

Figure 1 of the paper, redrawn: a retriever (query encoder plus document index) feeds the top-K passages to a generator, and the outputs are combined across passages. Based on Lewis et al. (2020), Figure 1.

Reading it: the question x enters on the left. The query encoder turns it into a vector, and MIPS compares that vector against every passage vector in the index to pick the top K. The dashed box on the left is the non-parametric memory: plain text you can read and edit. The generator on the right runs once for each retrieved passage, reading the question and that passage together. The final Σ box combines the K drafts into one answer, weighting each by how relevant its passage looked. That weighting is the paper's central trick, spelled out in §2.1.

2.1 Two ways to mix documents: RAG-Sequence and RAG-Token · original

“In one approach, RAG-Sequence, the model uses the same document to predict each target token. The second approach, RAG-Token, can predict each target token based on a different document.”Lewis et al. (2020), §2

Everyday picture

You ask two friends to answer a quiz question using a stack of reference cards. The first friend (RAG-Sequence) picks one card and writes the whole answer from it, and you average over which card they might have picked. The second friend (RAG-Token) may glance at a different card for every word. When an answer needs two facts from two cards, the second friend has the edge.

Tiny example (two passages, a two-word answer)

The retriever is 70% sure of passage z₁ and 30% sure of z₂. The answer is two tokens, y₁ then y₂. Given z₁, the generator gives y₁ probability 0.9 but y₂ only 0.2 (z₁ covers the first half of the answer). Given z₂, it gives y₁ 0.1 and y₂ 0.8 (z₂ covers the second half).

  • RAG-Sequence: commit to one passage for the whole answer, then average: 0.7 × (0.9 × 0.2) + 0.3 × (0.1 × 0.8) = 0.126 + 0.024 = 0.150.
  • RAG-Token: average per token, then multiply: (0.7 × 0.9 + 0.3 × 0.1) × (0.7 × 0.2 + 0.3 × 0.8) = 0.66 × 0.38 = 0.2508.

RAG-Token rates this answer higher because it can take the first word from z₁ and the second from z₂.

The math

In words: “RAG-Sequence scores the whole answer once per passage and then averages the scores, weighted by how relevant each passage looked. RAG-Token averages over passages separately for every word, then multiplies the words' averaged probabilities together.”

With the numbers: the only difference is whether Σ sits outside Π (Sequence: 0.150) or inside it (Token: 0.66 × 0.38 = 0.2508).

In Python:

import math
# p_η(z | x) for passages z₁, z₂
p_eta = [0.7, 0.3]
# p_θ(y_i | x, z, y_1:i−1): a row per passage, a column per token
p_theta = [[0.9, 0.2], [0.1, 0.8]]
# Σ outside Π
seq = sum(p_z * math.prod(p_y) for p_z, p_y in zip(p_eta, p_theta))
round(seq, 3)  # → 0.15
# Π outside Σ
tok = math.prod(sum(p_eta[z] * p_theta[z][i] for z in range(2)) for i in range(2))
round(tok, 4)  # → 0.2508

Try it: move the probabilities

Reading it: the top slider sets how much the retriever trusts passage z₁ (z₂ gets the rest). The other four sliders set how likely the generator finds each answer word when reading each passage. The two bars show the probability each method gives the full answer. Set things up so each passage supports a different word (the starting position), and RAG-Token wins. Make one passage support both words, and the two methods converge. That is exactly the paper's finding: RAG-Token did better on Jeopardy question writing, where one clue often combines two facts from two passages.

Why it matters today

Modern RAG systems rarely marginalise like this. They paste the top passages into one prompt and let a large model read them all at once, which behaves more like RAG-Token (any word can draw on any passage). The idea of weighting evidence by retrieval confidence survives in rerankers and in citation scoring. See the RAG lesson.

2.2 The retriever: Dense Passage Retrieval · original

Everyday picture

Every passage in the library gets a map coordinate, and so does your question. Passages that answer the question sit close to it on the map. Finding them means finding the nearest points, which an index can do in milliseconds even with 21 million points. This is a bi-encoder: questions and passages are placed on the map separately, so all the passage coordinates can be computed once, ahead of time.

Tiny example

Suppose the question's vector dotted with passage z₁'s vector scores 2.0, and with z₂ scores 1.15. Exponentiate and normalise: e2.0 = 7.39 and e1.15 = 3.16, which total 10.55. So p(z₁ | x) = 7.39 / 10.55 = 0.70 and p(z₂ | x) = 0.30: the retrieval weights used in §2.1.

In words: “encode the passage and the question separately into vectors; the more their dot product, the more probable the passage; normalise over the retrieved passages so the probabilities add to 1.”

With the numbers: scores (2.0, 1.15) → exp → (7.39, 3.16) → divide by 10.55 → (0.70, 0.30). That is softmax over the retrieved passages.

In Python:

import math
# d(z)ᵀ q(x) for passages z₁, z₂
scores = [2.0, 1.15]
e = [math.exp(s) for s in scores]
[round(v, 2) for v in e], round(sum(e), 2)  # → ([7.39, 3.16], 10.55)
# normalise: softmax
[round(v / sum(e), 2) for v in e]  # → [0.7, 0.3]

Why it matters today

The bi-encoder plus nearest-neighbour index is still the first stage of nearly every retrieval system. See the DPR companion, the retrieval lesson and the vector index lesson.

2.3 The generator: BART · original

Everyday picture

The writer is BART, a 400-million-parameter encoder-decoder transformer that is good at rewriting text. To show it a passage, the paper simply glues the question and the passage together into one input. No special machinery is needed: reading two texts side by side is just reading one longer text.

Tiny example

Input to the generator for passage z₂: “define middle ear” followed by the text of z₂. Output: an answer sentence, one token at a time, each with a probability. With K = 5 passages, the generator runs 5 times per question.

Why it matters today

Gluing the evidence into the input is exactly how modern RAG prompts work, except that today the “generator” is a much larger decoder-only chat model and the passages are wrapped in labelled blocks so the model can cite them. See the context engineering lesson.

2.4 Training · original

Everyday picture

Nobody tells the system which passage was the right one to fetch. It only sees question-answer pairs. If fetching a passage makes the correct answer more likely, training nudges the query encoder to fetch passages like it again. The retriever learns what is useful from the writer's success, like an assistant who learns which cards the writer actually uses.

In words: “for each training pair, take minus the log of the probability the whole system gives the correct answer, after mixing over passages; add these up and make the total small.”

With the numbers: if RAG-Token gives the correct answer probability 0.2508, that pair contributes −ln 0.2508 = 1.38 to the loss. Raising the probability to 0.5 would cut it to 0.69.

In Python:

import math
# p(y_j | x_j) for each training pair j
p = [0.2508]
round(sum(-math.log(p_j) for p_j in p), 2)  # → 1.38
round(-math.log(0.5), 2)  # → 0.69

One practical shortcut: the passage encoder and the 21-million-vector index are frozen. Only the query encoder and BART are fine-tuned. Re-encoding the whole index after every update would be very expensive, and the paper found it unnecessary.

Why it matters today

Most production systems freeze the retriever entirely and never train it end to end; they improve retrieval by choosing better embedding models, hybrid search and rerankers instead. The loss here is ordinary cross-entropy, covered in the loss functions lesson.

2.5 Decoding · original

Everyday picture

RAG-Token can write word by word with the per-word mix, so it plugs straight into a standard beam search. RAG-Sequence cannot: its probability only makes sense for complete answers. So it writes a few candidate answers per passage, then scores every candidate under every passage and adds up.

  • Thorough decoding: rescore every candidate under every passage, with extra model runs for any candidate a passage did not produce itself. Exact, but slow for long outputs.
  • Fast decoding: assume a passage gives probability ≈ 0 to any candidate it did not produce. No extra runs.

Why it matters today

Chat models today sample one answer from one prompt that holds all the passages, so this complexity disappears. The trade-off it illustrates (exact scoring versus a cheap approximation) comes up constantly in inference.

3 Experiments · original

The setup, in numbers

  • Library: the December 2018 Wikipedia dump, split into disjoint 100-word chunks: 21 million passages.
  • Index: FAISS with an HNSW approximate nearest-neighbour index.
  • Retrieved per question: K = 5 or 10 during training.
  • Size: 626 million parameters as the paper counts them (two 110M-parameter BERT encoders plus 406M for BART). Because the passage encoder stays frozen, 110M + 406M = 516 million of them are actually trained. The index is 21 million vectors, about 15.3 billion numbers.

Four kinds of task

TaskExampleScored by
Open-domain question answering (Natural Questions, TriviaQA, WebQuestions, CuratedTrec)“Who wrote …?” with no passage suppliedexact match
Abstractive QA (MS-MARCO)a full-sentence answerBLEU, ROUGE-L
Jeopardy question generationgiven “The World Cup”, write a clue that has it as the answerQ-BLEU-1 and human judges
Fact verification (FEVER)is this claim supported, refuted, or unverifiable?accuracy

Why it matters today

“Split the source into roughly 100-word passages and index their embeddings” is still the default starting point. How you chunk often matters more than which embedding model you pick; see the retrieval lesson.

4 Results · original

Everyday picture

On trivia-style questions, the open-book student beat both the biggest closed-book student (a model 18 times larger) and the specialist who only highlights answers inside the fetched text.

Selected numbers from Lewis et al. (2020), Table 1: Natural Questions, exact match (higher is better). Summarised with attribution.
SystemHow it answersNQ exact match
T5-11B (11 billion parameters)closed book34.5
REALMretrieve, then extract a span40.4
DPRretrieve, rerank, extract a span41.5
RAG-Tokenretrieve, then generate44.1
RAG-Sequenceretrieve, then generate44.5

Reading it: each bar is one system's exact-match score on Natural Questions, on an axis from 0 to 50. The closed-book T5 is the shortest bar despite having about 18 times RAG's 626 million parameters: memorising more is a poor substitute for looking things up. The two retrieve-then-extract systems are close to RAG, but RAG gets there without a separate reranker or span extractor. It also answered 11.8% of questions correctly even when no retrieved passage contained the answer, which an extractor, limited to copying text, can never do.

More factual and more specific (human judges)

On Jeopardy clue writing, judges compared 452 pairs of clues from BART alone and from RAG-Token. BART was judged more factual in 7.1% of pairs; RAG in 42.7%. RAG was also judged more specific by a wide margin.

Did it fetch the right evidence?

On FEVER, the top retrieved passage came from a correct evidence article in 71% of cases, and a correct article was in the top 10 in 90% of cases, even though the retriever was never told which evidence was correct.

Why it matters today

The habit worth copying is measuring retrieval separately from generation. If the right passage is not retrieved, no amount of prompt tuning fixes the answer. That is recall@k, taught in the metrics lesson.

Step through one question

This walks a single question through the whole pipeline with numbers small enough to follow. The vectors and probabilities are illustrative, chosen by hand to show the mechanics; they are not taken from the paper's model.

Reading it: press Next step to move through the five stages from Figure 1. Watch two things. First, the retriever never reads the answer; it only compares vectors. Second, the passage that looks most similar is not the only one that counts: the final answer's probability is a weighted mix across all the retrieved passages, so a lower-ranked passage that supports the answer still helps.

4.5 Hot-swapping the index · original

“This shows we can update RAG’s world knowledge by simply replacing its non-parametric memory.”Lewis et al. (2020), §4.5

Everyday picture

Replace an encyclopedia on the shelf with a newer edition and a student who looks things up gives newer answers the same afternoon. A student who memorised the old edition needs to be retaught.

Index: Questions about:

Reading it: the paper asked “Who is {position}?” for 82 world leaders who changed between December 2016 and December 2018. Pick an index and which year's office holders count as correct. When the two match, the same trained model is right about 70% of the time; when they do not match, it drops to 4 to 12%. Nothing about the model changed between the runs: only the library.

Why it matters today

This is the main reason businesses choose retrieval over fine-tuning for knowledge: update a document and the next answer reflects it. The flip side is operational: the index must be kept in sync with the source, including deletions and permission changes. See the training stages lesson for fine-tune versus RAG.

Also in §4.5

  • Learned retrieval matters. Freezing the retriever hurt every task. Swapping in keyword search (BM25) did best only on FEVER, whose claims are heavy with names, which is exactly where keyword matching shines.
  • More passages at test time. Retrieving more kept improving RAG-Sequence on Natural Questions, while RAG-Token peaked at 10.
  • Retrieval collapse (Appendix H). On some tasks, such as story writing, the retriever learned to fetch the same passages whatever the input, and the generator learned to ignore them. The result behaves exactly like BART alone.

5–6 Related work and discussion · original

The paper positions RAG as the general-purpose version of many single-task retrieval systems (question answering, fact checking, dialogue), and contrasts its readable, editable text memory with memory networks that store opaque vectors. It closes by suggesting that retriever and generator might one day be pre-trained together from scratch, and flags the usual risks of any fluent text generator, plus one specific to RAG: the library itself can be biased or wrong, and the model will faithfully repeat it.

What changed since 2020

The core, retrieve relevant text and condition generation on it, is unchanged. Almost everything around it has been re-engineered:

In the paperCommon todayWhyLearn it
Retriever and generator fine-tuned togetherFrozen general-purpose LLM, retrieval improved separatelyLarge models read retrieved text well without trainingRAG
Dense retrieval onlyHybrid search (dense + BM25) plus a cross-encoder rerankerExact IDs, codes and names need keyword matching; reranking sharpens the topretrieval
Marginalise over K generator runsOne prompt holding the top passages in labelled blocksOne model call instead of K; the model can compare passagescontext
Any passage for any userPermission-aware retrieval (filter by access rights before generation)Enterprise documents have owners and access listsRAG
The model answers; passages are implicitAnswers cite their sources, and citations are checkedUsers need to verify; grounding can be measuredevals
Always retrieve onceAgentic RAG: the model decides whether and what to search, and can search againMulti-part questions need several lookupsagent loop

Two later companions refine the retrieval step: HyDE searches with a hypothetical answer instead of the question, and Lost in the Middle shows where in the prompt retrieved passages should go.

Glossary

Every term with hover guidance on this page, in one place.