Lost in the Middle, annotated
How to read this page
- Any dotted word explains itself when you hover, tab to, or tap it, and so does every symbol in every equation.
- The position slider uses the paper's measured numbers: move the relevant document through the context and watch accuracy fall and recover.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The code version is the context assembler in the context engineering lesson.
Abstract
“Performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts.”Liu et al. (2023), Abstract. Read the original
Everyday picture
Read someone a list of twenty grocery items and ask what they remember. Most people recall the first few and the last few; the middle blurs. Psychologists call this the serial-position effect. This paper found the same shape in large language models: give a model twenty documents, one of which holds the answer, and it finds the answer far more often when that document is first or last than when it is in the middle.
What the paper claims
- Accuracy follows a U-shaped curve in the position of the relevant information: high at the start, lowest in the middle, partly recovered at the end.
- Models advertised with longer context windows are not better at using the context they already had.
- Even the simplest possible retrieval task (look up a random key and return its value) shows the effect for several models.
- Adding more retrieved documents quickly stops helping, long before the retriever stops finding more answers.
Why it matters today
Every RAG system and long-running agent decides what order to put things in the prompt. This paper is why careful systems put the most relevant material first (or last), keep the number of passages small, and test how accuracy changes with position instead of trusting the advertised context length.
1 Introduction · original
Everyday picture
A bigger desk lets you spread out more papers, but it does not mean you read them all equally carefully. Context windows had grown from 2,000 tokens to 100,000, and the question nobody had carefully measured was: does the model actually use everything on the desk?
The experimental idea
Keep the task fixed and change only two things: how much is in the context (10, 20 or 30 documents), and where the one useful document sits. If a model used its context robustly, moving the useful document would not matter. It does matter, a lot.
Tiny example
The same question, the same 20 documents, only the order changes. For GPT-3.5-Turbo, answer accuracy is 75.8% when the useful document comes first and 53.8% when it is 10th. The closed-book score, with no documents at all, is 56.1%. So in the worst position, giving the model the answer hurt compared with giving it nothing.
2 Multi-document question answering · original
2.1 The experiment · original
Everyday picture
Hide one page containing the answer in a stack of pages that look relevant but do not contain it, then ask the question. The distractors are hard on purpose: they were fetched by a real search system as the most relevant-looking passages that nonetheless lack the answer.
- Questions: 2,655 real search queries from NaturalQuestions-Open whose answer is a Wikipedia paragraph.
- Documents: Wikipedia chunks of at most 100 tokens: one with the answer, and k − 1 distractors retrieved by Contriever that do not contain it, ordered by decreasing relevance.
- Scoring: accuracy, meaning whether any correct answer string appears in the output.
- Decoding: greedy, so the results are repeatable.
- Models: MPT-30B-Instruct, LongChat-13B (16K), GPT-3.5-Turbo (4K and 16K) and Claude-1.3 (8K and 100K).
Hover or tap a block. The green document is the one that matters.
Reading it: the prompt is read left to right: an instruction, then the documents, then the question, and the model writes its answer last. The experiment changes only one thing between runs: which slot the green document occupies. Moving it changes nothing about what the right answer is, so any change in accuracy is caused purely by position.
2.3 The U-shaped curve · original
“These results indicate that current models cannot effectively reason over their entire context window when prompted for downstream tasks.”Liu et al. (2023), §2.3
Try it: move the useful document
Reading it: the strip of boxes is the context, one box per document, with the useful one highlighted. The chart's x-axis is the useful document's position (1 = first) and the y-axis is answer accuracy. The solid line joins the positions the paper actually measured (1, 5, 10, 15 and so on); the flat grey line is the same model's closed-book score, with no documents at all, and the orange vertical line is where the slider has put the useful document. Slide the position: accuracy drops steeply after position 1, bottoms out in the middle, and partly recovers at the end. For GPT-3.5-Turbo with 20 or 30 documents, the middle of the curve falls below the closed-book line. Readouts at unmeasured positions are linearly interpolated and labelled as such.
Two numbers that summarise a curve
In words: “the gap is the best position's accuracy minus the worst position's; the average is what you get if the useful document lands in a random measured position.”
With the numbers: GPT-3.5-Turbo, 20 documents, measured at positions (1, 5, 10, 15, 20): (75.8, 57.2, 53.8, 55.4, 63.2). The gap is 75.8 − 53.8 = 22.0 points. The average is 305.4 / 5 = 61.1%. Putting the best document first (75.8%) is worth almost 15 points over leaving its position to chance.
In Python:
# A(p): accuracy at position p
A = {1: 75.8, 5: 57.2, 10: 53.8, 15: 55.4, 20: 63.2}
# Δ: best minus worst
round(max(A.values()) - min(A.values()), 1) # → 22.0
# Ā: average over P
round(sum(A.values()) / len(A), 1) # → 61.1
The paper proposes exactly this kind of test as a standard: to claim a model can use a long context, show that its best-case and worst-case positions score about the same.
Longer windows do not fix it
Where a prompt fits in both a model and its extended-context sibling, their curves nearly coincide: GPT-3.5-Turbo and its 16K version, and Claude-1.3 and its 100K version, score within a few tenths of a point of each other at every position. A bigger window lets more text in; it does not make the model read the middle more carefully.
Why it matters today
Newer models have improved on these tests, but the lesson is a method rather than a number: measure your own model on your own prompt with the key evidence placed at several positions. The context engineering lesson builds a prompt assembler that puts the highest-priority material at the edges.
3 Key-value retrieval: the simplest possible test · original
Everyday picture
If a model struggles to use a document in the middle, is the problem understanding, or merely finding? To separate the two, the authors built a test with no meaning at all: a long list of random ID pairs, and a request to return the partner of one ID. It is pure lookup, like Ctrl+F.
Tiny example
Reading it: this is the shape of the real task at toy size. Every key and value is a random 128-bit identifier written as text (a UUID), so there is no meaning to exploit: the model must match the key character for character and copy the value after it. The highlighted pair is the target, and the slider moves it through the list. The paper used 75, 140 or 300 pairs, with 500 examples each.
What happened
- Claude-1.3 and its 100K version were nearly perfect at every length and position.
- GPT-3.5-Turbo (both versions) and MPT-30B-Instruct were worst when the target pair was in the middle, especially with 140 or 300 pairs.
- The worst case across models was 45.6% accuracy, on a task that is only exact copying.
Why it matters today
Synthetic lookup tests of this kind, often called “needle in a haystack” tests, now appear in many long-context model reports. They are a floor, not a ceiling: passing a pure-lookup test does not mean a model will reason well over the middle of a long input, which is what §2 measures.
4 Why does it happen? · original
The paper runs three preliminary investigations. None fully explains the effect, but each narrows it down.
4.1 Model architecture · original
Everyday picture
A decoder-only model reads strictly left to right, so while reading document 3 it cannot know what question is coming. An encoder-decoder model reads the whole input at once, in both directions, before answering, so every document is read in the light of every other document.
What they found
Encoder-decoder models (Flan-UL2, Flan-T5-XXL) were fairly robust within the input lengths they were trained on: Flan-UL2 varied by only 1.9 points between its best and worst positions with inputs of up to 2,048 tokens. Past their training length, they too developed the U shape.
4.2 Query-aware contextualization · original
Everyday picture
If a left-to-right reader cannot know the question while reading the documents, tell them the question first. The authors put the question both before and after the documents.
Hover or tap a block.
Reading it: in the top row the question comes last, so while the model's layers process each document, the question is invisible to them (a decoder's causal mask hides everything to the right). In the bottom row a copy of the question comes first, so every document is processed knowing what to look for. The copy at the end is kept so the question is fresh when the answer is written.
What they found
- Key-value lookup: near-perfect for every model at every length. GPT-3.5-Turbo (16K) went to 100% with 300 pairs.
- Question answering: barely changed. A small gain when the useful document was first, a small loss elsewhere.
So for pure lookup, knowing the key in advance fixes the problem; for reasoning over documents, it does not.
Why it matters today
Stating the task before a long block of data, and repeating it after, is cheap and still widely recommended, especially for extraction and lookup tasks. See the context engineering lesson.
4.3 Instruction fine-tuning · original
Everyday picture
Instruction-tuning data usually puts the instruction at the very start. Perhaps that trains models to over-weight the beginning? To test it, the authors compared MPT-30B-Instruct with its base model, MPT-30B, before any instruction tuning.
What they found
- Both show the U shape, so instruction tuning is not the cause. Tuning actually narrowed the best-to-worst gap, from nearly 10 points to about 4.
- In Llama-2 models (Appendix E), the 7B model showed only a recency bias (the end is favoured); the U shape appeared only in the 13B and 70B models.
5 Is more context always better? · original
Everyday picture
Handing a reader twice as many pages makes it more likely the answer is somewhere in the pile, and less likely they find it. The authors ran a realistic RAG setup, retrieving the top k Wikipedia passages for real questions, and tracked two curves as k grew: how often the retriever had fetched an answer (recall), and how often the model actually answered correctly.
What they found
The model's accuracy flattened long before the retriever's recall did. Going from 20 to 50 retrieved passages improved accuracy by only about 1.5 points for GPT-3.5-Turbo and about 1 point for Claude-1.3, while making every prompt much longer, slower and more expensive.
Why it matters today
Two practical fixes follow directly, and the paper names both: rerank so the most relevant passages sit at the start, and truncate the list so fewer, better passages go in. Both are implemented in the RAG lesson; the cost side is in the cost lesson.
6–7 Related work and conclusion · original
The authors connect their U-shaped curve to the serial-position effect from psychology: in free recall of a list, people best remember the first items (primacy) and the last items (recency). Seeing the same shape in a transformer is surprising, since self-attention can in principle look at any position equally easily. The conclusion restates the main finding and offers the position-sweep protocol as a new way to evaluate long-context models.
Using it today
| Finding | What careful systems do | Learn it |
|---|---|---|
| The start and end are used best | Put the highest-ranked passages first; put instructions and the question where they will be seen | context |
| More passages quickly stop helping | Rerank, then send only the top few; measure accuracy against k instead of guessing | RAG |
| Advertised context length is not usable context | Test your own prompts with the key fact at several positions | evals |
| Long histories bury early details | Summarise or drop old turns in long conversations and agent runs, the problem known as context rot | memory |
Retrieval itself is covered by the RAG companion and the HyDE companion.
Glossary
Every term with hover guidance on this page, in one place.