HyDE: search with a made-up answer, annotated
How to read this page
- Any dotted word explains itself when you hover, tab to, or tap it. So does every symbol in every equation.
- The embedding map is the heart of the page: drag nothing, just toggle and watch which real document wins.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The code version is hyde_search in the RAG lesson.
Abstract
“The generated document is not real, can contain factual errors but is like a relevant document.”Gao et al. (2022), §1. Read the original
Everyday picture
You want a recipe but can only describe the dish badly: “that orange soup with coconut, from somewhere in Asia?” Searching with that question finds other people asking similar questions. Instead, you scribble what you think the recipe looks like (“Thai pumpkin soup: simmer pumpkin with coconut milk, red curry paste and lime…”), and hold that against the cookbook. Your guess may be wrong in details, but it looks like a recipe, so it lands on the right page. HyDE does this with a language model writing the guess and an embedding model doing the matching.
What the paper claims
- Strong search with no relevance labels at all: no example of which document answers which question is needed.
- It beats the same embedding model used the normal way on every English benchmark tested, and is competitive with retrievers fine-tuned on large labelled datasets.
- No model is trained. An off-the-shelf instruction-following model and an off-the-shelf embedding model are simply chained.
Why it matters today
“Rewrite the query before searching” is now a standard move in RAG systems, and HyDE is its most famous form. It is also a clean lesson in how questions and answers live in different neighbourhoods of an embedding space.
1 Introduction · original
Everyday picture
A dense retriever is trained by showing it thousands of (question, correct passage) pairs until questions land near their answers. Where do those pairs come from? Mostly one big dataset, MS-MARCO, which restricts commercial use and does not cover your legal contracts or your support tickets. Without such pairs you are doing zero-shot retrieval: the system must work on data it has never been tuned for.
The paper's move
Split search into two jobs that each need no labels:
- Generate: an instruction-following model writes a passage that would answer the question (the “hypothetical document”).
- Match: an unsupervised embedding model, trained only to put similar documents near each other, finds real documents near the hypothetical one.
The question is never compared with a document directly. Relevance is handled by the writer; similarity by the embedder.
Why it matters today
Most organisations start a search project with zero labelled pairs. HyDE is a way to get strong retrieval on day one, before any search logs exist.
2 Related work · original
The paper sits between three lines of work: dense retrieval (embed questions and documents, search by dot product, as in DPR), instruction-following models that generalise to new tasks from a written instruction (InstructGPT), and zero-shot retrieval benchmarks such as BEIR. It also contrasts itself with generative retrieval, which trains a model to output document IDs; HyDE needs no training and uses an ordinary vector index.
3 Method · original
Hover or tap any block in the redrawn figure.
Hover or tap a block, starting with query q on the left.
Reading it: read left to right. The query goes to the instruction-following model with an instruction such as “write a passage to answer the question”, and the model writes N short fake passages. The encoder turns each fake passage into a vector, and the vectors are averaged. That average, not the query's own vector, is compared against the precomputed vectors of the real documents (the box below), and the nearest real documents are returned. The dashed wire is the variant in Equation 8, where the query's own vector joins the average as one more vote.
3.1 Dense retrieval, decoded · original
Everyday picture
Normal dense search needs two translators into one shared map language: one for questions, one for documents. Teaching them to agree on where a question and its answer should land takes labelled examples, which is exactly what we do not have.
In words: “turn the question into a vector with one encoder, the document into a vector with another, and score their similarity by the inner (dot) product.”
With the numbers: with vq = (0.2, 0.8) and a document at (0.85, 0.35), sim = 0.2 × 0.85 + 0.8 × 0.35 = 0.17 + 0.28 = 0.45.
In Python:
# enc_q(q)
v_q = (0.2, 0.8)
# enc_d(d)
v_d = (0.85, 0.35)
# ⟨v_q, v_d⟩
sim = sum(a * b for a, b in zip(v_q, v_d))
round(sim, 2) # → 0.45
Why it matters
The hard part is making encq and encd agree. HyDE sidesteps it by using only the document encoder: a question becomes a document first.
3.2 HyDE · original
“Here, we expect the encoder’s dense bottleneck to serve a lossy compressor, where the extra (hallucinated) details are filtered out from the embedding.”Gao et al. (2022), §1
Everyday picture
Ask five friends to each write a one-paragraph answer from memory. Each gets some details wrong, and in different ways. Average their paragraphs into one “typical answer” and the idiosyncratic mistakes cancel, leaving the shape of a good answer. The embedding step helps too: squeezing a paragraph into a few hundred numbers keeps its gist (topic, kind of text) and drops specifics such as the exact wrong year.
Tiny example (two-dimensional vectors)
The model writes N = 2 fake passages, encoded as f(d̂₁) = (0.9, 0.1) and f(d̂₂) = (0.7, 0.5). Their average is ((0.9 + 0.7)/2, (0.1 + 0.5)/2) = (0.8, 0.3). The real relevant passage sits at (0.85, 0.35); a real forum post that merely asks the same question sits at (0.3, 0.9); the original query sits at (0.2, 0.8).
| Search with | score vs relevant passage | score vs forum post | Winner |
|---|---|---|---|
| the query (0.2, 0.8) | 0.2·0.85 + 0.8·0.35 = 0.45 | 0.2·0.3 + 0.8·0.9 = 0.78 | forum post (wrong) |
| HyDE average (0.8, 0.3) | 0.8·0.85 + 0.3·0.35 = 0.785 | 0.8·0.3 + 0.3·0.9 = 0.51 | relevant passage |
The question looks like other questions; the fake answer looks like real answers.
The math
In words: “ask the instruction model to write N passages answering the question, embed each with the document encoder, and average the vectors (optionally including the question's own vector); search with that average.”
With the numbers: N = 2 gives (0.8, 0.3). Counting the query too gives ((0.9 + 0.7 + 0.2)/3, (0.1 + 0.5 + 0.8)/3) = (0.6, 0.467), which still prefers the relevant passage (0.673 against 0.600), though by less.
In Python:
# f(d̂_k) for the N = 2 fake passages
f_d = [(0.9, 0.1), (0.7, 0.5)]
# f(q)
f_q = (0.2, 0.8)
N = len(f_d)
# (1/N) Σ_k f(d̂_k)
v_hat = [sum(v[i] for v in f_d) / N for i in range(2)]
[round(x, 3) for x in v_hat] # → [0.8, 0.3]
# query as one more vote
v_hat = [(sum(v[i] for v in f_d) + f_q[i]) / (N + 1) for i in range(2)]
[round(x, 3) for x in v_hat] # → [0.6, 0.467]
relevant, forum = (0.85, 0.35), (0.3, 0.9)
[round(sum(a * b for a, b in zip(v_hat, doc)), 3) for doc in (relevant, forum)] # → [0.673, 0.6]
The instructions are short and task-specific. For web search the paper uses “Please write a passage to answer the question”; for scientific claims, “Please write a scientific paper passage to support/refute the claim” (Appendix A.1).
Why it matters today
The averaged vector is the paper's estimate of the expected embedding of a good answer. The same “embed several samples and average” trick is used to make embeddings of noisy or short inputs more stable.
Try it: the embedding map
A two-dimensional stand-in for the real embedding space (which has 768 dimensions). Positions are illustrative and chosen so the arithmetic matches the tiny example above; the ranking is computed live.
Reading it: grey squares are real documents, blue dots are the fake passages the model wrote, the red diamond is the query, and the star is the vector HyDE actually searches with. The line under the map ranks the real documents by inner product with the star. At N = 0 the star is the query, and the forum post that asks the same question wins. Slide N up and the star moves into the “answers” region, where the relevant passage wins. Tick the checkbox and the query pulls the star back towards the questions: with many fake passages it barely matters; with one it can.
4 Experiments · original
Setup
- Writer: InstructGPT (text-davinci-003), sampled at temperature 0.7.
- Encoder: Contriever, trained with unsupervised contrastive learning (no relevance labels); mContriever for other languages.
- Benchmarks: web search (TREC DL19 and DL20), six low-resource BEIR tasks, and four languages from Mr.TyDi.
- Metric: mainly nDCG@10 (how good the top 10 results are, rewarding relevant results higher up) and recall.
Web search
| System | Relevance labels used? | DL19 | DL20 |
|---|---|---|---|
| BM25 (keyword search) | no | 50.6 | 48.0 |
| Contriever (the same encoder, used normally) | no | 44.5 | 42.1 |
| HyDE (InstructGPT + Contriever) | no | 61.3 | 57.9 |
| Contriever fine-tuned on MS-MARCO | yes, hundreds of thousands | 62.1 | 63.2 |
Reading it: each pair of bars is one system on DL19 and DL20, on an axis from 0 to 70. Unsupervised Contriever used the normal way is worse than keyword search. The same encoder searching with HyDE's fake answers jumps past BM25 and nearly reaches the version fine-tuned on a large labelled dataset, which is effectively an upper bound here.
Low-resource and multilingual search
On six BEIR tasks, HyDE improved Contriever everywhere. The most dramatic: TREC-COVID nDCG@10 went from 27.3 to 59.3 (BM25: 59.5). In other languages the gains were smaller but real, for example Japanese MRR@100 from 19.5 to 30.7, and the paper notes both the small multilingual encoder and the writer are weaker outside English.
Why it matters today
Two lessons carry over. Always compare against BM25: a “semantic” retriever can lose to keyword search on a new domain. And measure retrieval on its own with recall@k or nDCG before judging the whole RAG system; see the metrics lesson.
5 Analysis · original
A better writer gives better search
Reading it: each bar is HyDE's DL19 nDCG@10 with a different instruction-following model writing the fake passages, from 11 billion to 175 billion parameters (numbers from Gao et al., 2022, Table 4). All three beat plain Contriever (44.5, the grey bar), and bigger writers help more. The writer is doing the “relevance” work, so a smarter writer produces fake answers that land closer to real ones.
With an already fine-tuned encoder
HyDE is meant for when you have no labels, but the paper also tried it on top of the fine-tuned Contriever. With the strongest writer it still helped (DL19: 62.1 → 67.4); with smaller writers it slightly hurt. So the fake answers carry some signal even a supervised retriever misses, but only when they are good.
6 Conclusion · original
The authors raise a provocative question: if a language model can capture relevance by writing an example answer, is a learned numerical relevance score even necessary? They leave it open, and end with a practical lifecycle. On day one of a new search system, serve queries with HyDE. As search logs accumulate, train a supervised retriever and route more traffic to it, keeping HyDE for rare and emerging queries.
Using it today
| Consideration | What to do |
|---|---|
| Cost and latency: one extra model call per search | Use a small, fast model for the fake passage; cache results for repeated queries |
| The fake answer can encode a misconception and steer search towards it | Combine HyDE with the plain query (Eq. 8) or with hybrid search, then rerank with a cross-encoder |
| Exact identifiers (error codes, part numbers) | A fake passage may invent a different code. Keep BM25 in the mix |
| Never show the fake passage to users | It is a search key, not evidence. Only real retrieved text should be cited |
HyDE is one of several query-rewriting tricks (multi-query, step-back questions, decomposition). All of them are implemented and compared in the RAG lesson; how the retrieved passages should then be arranged is the subject of the Lost in the Middle companion.
Glossary
Every term with hover guidance on this page, in one place.