primer.agents.rag

Retrieval-augmented generation (RAG), end to end

Run: python -m primer.agents.rag

This lesson builds on keyword search, embeddings and reranking from primer.ml.embeddings.retrieval, on prompt assembly from primer.agents.context, and on the tool-calling loop from primer.agents.agent_loop.

Level 1: The practitioner's guide

In one sentence. Retrieval-augmented generation (RAG) fetches the few passages of your own documents most likely to answer a question, puts them in the prompt, and has the model answer from those passages with citations, so the model can use knowledge it was never trained on and a reader can check where each fact came from.

When you need it. When the answer lives in documents the model has not seen (your policies, tickets, contracts, product manuals), when those documents change faster than you could retrain, when every answer must name its source, or when different users may read different documents. You don't need it when the model already knows the subject (public, stable knowledge), when you want to change how the model behaves rather than what it knows (that is fine-tuning: it teaches format and style, not facts that change), or when the whole collection fits in the prompt. Anthropic's contextual retrieval post puts that last threshold at about 200,000 tokens, roughly 500 pages: below it, put everything in the prompt and skip the pipeline. The tell: your model answers a question about the 2023 travel policy confidently and wrongly, because it never saw the 2026 one.

Your options. From the cheapest to the most capable, each one usually added on top of the last:

Option What it does What it guarantees What it costs Where it lives
Everything in the prompt Sends the whole collection with every question Nothing is ever missed by a search Every token, every call; only up to a few hundred pages Your prompt
Keyword search (BM25) Ranks passages by the question's rarer words Exact identifiers (ERR-4012, ticket numbers) are found An inverted index; misses synonyms and paraphrase A search engine
Dense search Ranks passages by embedding similarity, meaning over words Paraphrases are found ("scam message" finds the phishing guide) An embedding model at ingest and per query, a vector index; blurs exact codes A vector index
Hybrid search Runs both and fuses the rankings by rank (RRF), never by score Both kinds of question, and no score scales to reconcile Two indexes, two searches per question Both indexes plus a fusion step
Reranking A cross-encoder reads the question next to each of the top candidates Far better ordering of the top few One model score per candidate, so only over a shortlist A reranking model between search and the prompt
Retrieval upgrades Contextual chunks, query rewriting, HyDE, multi-query, date filters, parent-child Each fixes one named failure Model calls at ingest (context lines) or per query (rewrites, drafts) Ingestion or query time
Agentic RAG The model decides when, how often and with what words to search Multi-part and vague questions handled; no search for small talk Several model calls and a step budget per question An agent loop with a search tool
GraphRAG Extracts entities and relations first, answers by walking the graph or summarizing clusters Questions about connections and whole-collection themes A model pass over every document at ingest, and a graph to maintain Ingestion, plus a graph store

How to choose. Start from the questions people actually ask, with a labelled set of at least a dozen where you know the right document.

  • Questions full of exact identifiers and paraphrases both, which is what enterprise data looks like: hybrid search with a reranker. In this lesson's toy set, dense search alone never finds ERR-4012 and plateaus at 92% recall; hybrid reaches 100% by the second result; reranking the hybrid shortlist lifts the share of questions answered by the first result from 75% to 92%.
  • Chunks that lose their meaning when cut from the document (a fix that never names the error it fixes): contextual chunks. Prefixing each chunk with where it comes from turns "not in the top 20" into "second" here; Anthropic measured a 49% drop in top-20 retrieval failures from contextual embeddings and contextual BM25, and 67% with a reranker added.
  • Follow-up questions and vague phrasing: query rewriting and HyDE, which searches with a model-written draft answer.
  • Two questions in one, or a question the first search doesn't settle: agentic RAG.
  • "What depends on X?" and "what are the themes?": GraphRAG, and only then, because it is the most expensive to build and keep current.
  • Whatever you pick: filter by permissions and date before ranking, and verify citations before the answer reaches anyone. Retrieval quality is decided in the first boxes (parsing, chunking, search), never by the prompt.

What it costs. Ingestion costs an embedding per chunk, a keyword index, and for contextual chunks a model call per chunk: with prompt caching, Anthropic's post prices that at \$1.02 per million document tokens. A query costs one embedding, two searches, a reranker score for each of a handful of candidates (8 here, kept to 3), and one generation whose input is the question plus those passages; that prompt is what the answer's latency and token bill mostly consist of. Agentic RAG multiplies the model calls by the number of searches (two for the two-part question in this lesson, zero for "thanks, that's all!") and needs a step limit. Quality costs come from the front of the pipeline: a table flattened by a naive PDF extractor, so that the hotel cap floats free of its grade, is a failure no later stage can repair. Effort goes mostly into the labelled question set and the parsing, not the model.

What breaks.

  • The right passage is not in the index. Parsing garbled it, chunking split it from its context, or the permission sync dropped it. Check ingestion before touching anything downstream.
  • The right passage is in the index but not in the top k. Measure recall@k on the labelled set. If it is low, change retrieval (hybrid, reranker, rewriting, HyDE, contextual chunks); a prompt cannot fix it.
  • Permission leaks. Filtering after generation strips the citation and leaves the fact: in this lesson the restricted forecast's "41 million dollars" reaches an employee who may not read it. Copy each document's access list onto its chunks and filter before ranking; sync revocations promptly.
  • Stale sources. A superseded 2023 policy outranks the current one. Carry the date on every chunk and filter or prefer by it.
  • Confident answers with no support. Instruct the model to answer only from the sources and to decline otherwise, then check that every cited id was in the context and that its passage supports the claim.
  • Blaming the model. With dense-only retrieval this lesson's twelve questions produce five wrong answers, of which only one is a retrieval miss; switching to hybrid removes that one, and the two that remain are generation and data problems. Each needs a different fix, so diagnose first.

In the wild. The name comes from Lewis et al. (2020), who paired a retriever with a generator over a document index. Keyword search is BM25 in Lucene, Elasticsearch and OpenSearch; dense search runs on vector indexes such as FAISS and hosted services such as Pinecone; Elasticsearch documents RRF for fusing the two. HyDE is Gao et al. (2022). Contextual retrieval and prompt caching are described in Anthropic's post, and Claude's citations feature returns the exact passage behind each claim for PDF, plain-text and custom documents. GraphRAG is Microsoft's open implementation of Edge et al. (2024). RAGAS evaluates a RAG pipeline on retrieval and generation separately, which is exactly the split the debugging section relies on.

Go deeper. Level 2 builds the whole pipeline in plain Python: a PDF extractor's damage repaired step by step, chunks that carry their metadata, BM25 and dense search fused by a one-line formula, a reranker, a prompt that tags every source, a citation check, the permission leak reproduced, each upgrade as a working function with its before and after, and the diagnosis figure you can rerun on your own questions. If you only needed to choose, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

What follows builds every box of the pipeline from scratch, in the order a question travels through them, then the upgrades, then how to tell a retrieval failure from a generation failure.

An open-book exam with a librarian. Before you answer, the librarian fetches the few pages most likely to contain the answer; you answer from those pages and write down which page each fact came from. If the pages don't contain the answer, you say so instead of guessing.

Retrieval-augmented generation is that arrangement for a language model. The model's training knowledge is frozen and doesn't include your company's documents, so before each answer the system retrieves relevant passages from your documents and puts them in the prompt, and the model generates an answer from them, with citations. It's how you give a model knowledge that changes, or that must be cited, without retraining it.

flowchart LR subgraph ING["Ingestion, ahead of time"] D[Documents] --> P[Parse] --> C[Chunk + metadata] --> E[Embed + keyword index] end E --> IX[(Index)] subgraph QRY["Query time, per question"] Q[Question] --> F[Filter by permissions<br/>and date] --> R[Retrieve: hybrid] --> RR[Rerank] --> AC[Assemble context] --> G[Generate + cite] --> V[Verify citations] end IX --> F

Reading it: the top row runs once per document, whenever documents change; the bottom row runs for every question. Everything the answer can contain has to survive every box, left to right: a table mangled at Parse or a passage missing at Retrieve can't be fixed by a better prompt at Generate. That's why most RAG quality work happens in the first boxes, not the last.

In code: RAGIndex is the top row: it chunks every document and builds the keyword and dense indexes. answer is the bottom row, from filtered retrieval to verified citations, and returns a RAGResult.

The rest of this lesson walks the boxes in order, then covers the upgrades and how to debug a wrong answer.

Parsing: where quality dies first

Everyday picture. Photocopy a newspaper page and read it straight across: you get the first line of one column, then the first line of the next story, and nothing makes sense. Text extracted from PDFs has the same problems: running headers and page numbers mixed into the text, words broken across lines, and tables flattened into a row of numbers with no columns.

Worked example. MESSY_PDF is two pages of a travel policy as a naive extractor returns it. clean_extracted_text repairs it in four steps:

Damage Before After
Header repeated on every page Northwind Ltd Travel Policy 2026 … CONFIDENTIAL removed (it appears on 2 of 2 pages)
Page-number footer Page 1 of 2 removed
Word hyphenated across a line break all em- / ployees travelling all employees travelling
Table flattened L1-L3 150 60 L4-L6 200 75 kept as rows, then one sentence per row

The table is the dangerous one. Flattened, "200" floats free of its grade. table_rows_as_sentences turns each row into a self-contained sentence, Grade L4-L6: hotel cap 200, meal cap 75., so any row can be retrieved on its own and still says what its numbers mean.

flowchart LR RAW[Raw extracted text] --> H[Drop lines repeated<br/>on most pages] H --> N[Drop page numbers] N --> J[Rejoin hyphenated words,<br/>unwrap paragraphs] J --> T{Table rows?} T -->|yes| S[One sentence per row,<br/>headers repeated] T -->|no| P[Paragraphs] S --> OUT[Clean text for chunking] P --> OUT

Reading it: each box removes one kind of extraction damage. The diamond is the important decision: prose and tables need different handling, because a table's meaning lives in its columns. Real pipelines add layout-aware parsers and OCR (optical character recognition: reading text from an image of a page) for scanned documents, with a quality check on the output.

Why it matters. Enterprise documents are scanned, multi-column, full of tables that span pages. If parsing garbles the numbers, no retriever or model can recover them, and the failure looks like a model error.

In code: extract_table finds a table by its header row in the cleaned text and returns each row as a dict, ready for table_rows_as_sentences.

Chunking, with metadata that travels

Everyday picture. Index cards. Each card holds one idea, small enough to match a question precisely, and in its corner a label: which document, which date, who may read it.

Worked example. With a 15-word limit, it-004 splits into two passages:

Passage id Text Updated Readers
it-004#0 ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. 2026-04-10 everyone
it-004#1 Update AnyConnect to version 5.1 or later and reboot. 2026-04-10 everyone

(The limit is tiny because the toy documents are tiny; real systems use a few hundred tokens.) Split on the document's structure (sentences, paragraphs, sections), not at fixed character counts that cut a sentence in half. primer.ml.embeddings.retrieval compares chunking strategies in depth.

Why it matters. Small chunks match precisely but can lose context (the second card doesn't say which error it fixes; see contextual retrieval below). The metadata is what makes permission and date filters possible later.

In code: chunk_document splits a document on sentence boundaries into Passages, and each Passage carries its id, title, department, date and readers.

Retrieval: keyword and meaning, fused

Everyday picture. Two librarians: one searches the catalogue for your exact words, the other understands what you mean even when you use different words. You take the books both of them rank highly.

  • Keyword search (BM25) scores passages by the question's words, giving more weight to rare words and less to long passages. It nails exact identifiers like ERR-4012 and misses synonyms.
  • Dense search compares embeddings (vectors where closeness means similar meaning; see primer.ml.embeddings). It finds "scam message in my inbox" → the phishing guide with no shared words, but blurs exact codes.
  • Hybrid search runs both and merges the two rankings with reciprocal rank fusion (RRF), which only needs the ranks, never the scores. BM25 scores and cosine similarities live on different scales, so adding them would be meaningless.
Level 3: the formula and its symbols

$$ \text{RRF}(d) = \sum_{i} \frac{1}{k + \text{rank}_i(d)} $$

Symbols

Symbol Meaning here Worked example
$d$ one passage it-004#0
$i$ one of the rankings being fused (keyword, dense) 2 rankings
$\text{rank}_i(d)$ position of $d$ in ranking $i$ (1 = best); rankings that miss $d$ contribute nothing 1st, 3rd
$k$ a damping constant; 60 is the usual choice 60
$\sum_i$ add up over the rankings

In words: each ranking gives a passage a vote worth one over sixty-plus its position, and the votes are added.

On the worked example: a passage ranked 1st by dense search and 3rd by keyword search scores 1/61 + 1/63 = 0.0323. A passage ranked 1st by one list and absent from the other scores only 1/61 = 0.0164. Passages that both methods agree on rise to the top.

Level 3: in Python

In Python:

k = 60
# 1st by dense, 3rd by keyword: Σ_i 1/(k + rank_i(d))
round(1 / (k + 1) + 1 / (k + 3), 4)  # → 0.0323
# 1st in one list, absent from the other
round(1 / (k + 1), 4)  # → 0.0164

Dense search plateaus at 92% recall because it never finds the error code, while hybrid reaches 100% by k = 2 and reranking lifts its top-1 from 75% to 92%

Reading it: the x-axis is how many passages you keep (k); the y-axis is recall@k, the share of the 12 labelled questions whose correct document is among those k. On this small set many questions reuse the documents' own words, so keyword search is already strong. Dense search plateaus at 92% because it never finds ERR-4012. Hybrid reaches 100% by k = 2 but only 75% at k = 1. Reranking the hybrid shortlist lifts the top-1 result to 92%. Run this measurement on your questions before choosing a method.

Level 3: the formula and its symbols

$$ \text{recall@}k = \frac{1}{|Q|} \sum_{q \in Q} \mathbb{1}\bigl[\text{a relevant document is in the top } k \text{ for } q\bigr] $$

Symbols

Symbol Meaning here
$Q$ the labelled questions; $\lvert Q\rvert$ is how many (12 here)
$q$ one question
$\mathbb{1}[\ldots]$ 1 if true, 0 if not

In words: the share of questions for which the right document made it into the top k.

On the worked example: dense search at k = 3 finds the right document for 11 of 12 questions (all but ERR-4012): 11/12 = 0.92.

Level 3: in Python

In Python:

# 𝟙[...] for each of the 12 questions; ERR-4012 is the 0
found = [1] * 11 + [0]
# (1/|Q|) Σ over q in Q
round(sum(found) / len(found), 2)  # → 0.92

In code: RAGIndex.retrieve ranks the allowed passages with primer.ml.embeddings.retrieval.BM25, with dense similarity, or with both fused by primer.ml.embeddings.retrieval.reciprocal_rank_fusion, depending on the mode you ask for. recall_curve measures recall@k for each method.

Reranking: read the question and each passage together

Everyday picture. The librarians bring back twenty books quickly; an expert then reads your question next to each book and picks the best three.

A reranker (usually a cross-encoder: a model that reads the question and one passage together and outputs a relevance score) is far more precise than comparing precomputed vectors, and far too slow to run over every passage. So retrieval casts a wide, cheap net (here 8 candidates) and the reranker keeps the best few (here 3). This module reuses the toy cross-encoder from primer.ml.embeddings.retrieval.

In code: RAGIndex.rerank scores each (question, passage) pair with primer.ml.embeddings.retrieval.CrossEncoder and keeps the best few.

Assemble, generate, cite, verify

Worked example. "What does ERR-4012 mean?" becomes this context (each source escaped and tagged with its id; the question last; see primer.agents.context):

<sources>
<source id="it-004#1" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">Update AnyConnect to version 5.1 or later and reboot.</source>
<source id="it-004#0" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date.</source>
<source id="it-002#2" title="Password policy" updated="2025-11-20">This policy applies to all employees and contractors.</source>
</sources>
<question>What does ERR-4012 mean?</question>

and the answer is "ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. [it-004#0]". verify_citations then checks that every cited id was really in the context and that the cited passage supports the claim before it; a question the sources don't cover gets "I don't know based on the provided sources."

sequenceDiagram participant U as User participant A as RAG service participant I as Index participant M as Model U->>A: What does ERR-4012 mean? A->>I: hybrid search, filtered to what this user may read I-->>A: 8 candidate passages A->>A: rerank, keep 3 A->>M: rules + tagged sources + question M-->>A: answer with [it-004#0] A->>A: verify: id was in context, passage supports claim A-->>U: answer + clickable citation

Reading it: time runs downward. Note the order: the permission filter runs at the index, before any passage is fetched; the model only ever sees three passages; and nothing reaches the user until the citations check out. Citations are what let a person verify an answer in seconds, which is most of what makes people trust the system.

In code: assemble_context builds the tagged sources and the question. generate sends them to the model with the answer-from-sources rules, and grounded_policy is the offline stand-in model that honours them. answer chains retrieve, rerank, assemble, generate and verify_citations.

Permission-aware retrieval: filter before, never after

Everyday picture. A librarian checks your library card before fetching from the restricted archive. Checking it after you've read the document is pointless.

Each passage carries its access-control list (ACL: the groups allowed to read it), copied from the source system (SharePoint, Google Drive, a wiki). RAGIndex.retrieve drops every passage the user can't read before ranking, so restricted text can't reach the context, the model, or the answer.

Worked example. "What is the Q3 revenue forecast?" (the forecast document is readable only by finance and exec):

Who asks How it's filtered Answer
finance user before retrieval "Q3 revenue is forecast at 41 million dollars… [fin-006#0]"
anyone else before retrieval "I don't know based on the provided sources."
anyone else after generation (citation removed) "Q3 revenue is forecast at 41 million dollars…"
sequenceDiagram participant U as Employee (not finance) participant A as RAG service participant M as Model rect rgb(250, 225, 225) Note over A,M: Wrong: filter after generation A->>M: sources include the restricted forecast M-->>A: "…41 million dollars [fin-006#0]" A->>A: strip the restricted citation A-->>U: "…41 million dollars" (leaked) end rect rgb(225, 240, 225) Note over A,M: Right: filter before retrieval A->>A: drop passages the user can't read A->>M: sources without the forecast M-->>A: "I don't know…" A-->>U: nothing restricted end

Reading it: in the red box the model saw the restricted passage, so the number is already in its words, and deleting the citation marker afterwards leaves the fact behind. In the green box the passage never left the index. Permission changes in the source system must also sync to the index quickly, or a revoked user keeps access until the next re-index.

In code: post_generation_filter_answer is the wrong way, kept so the leak can be shown: it retrieves for every group, generates, then strips the restricted citation markers.

Retrieval upgrades

Each upgrade below is a working function, and each fixes a failure you can reproduce in demo():

Upgrade Everyday picture Before → after (top 3)
Query rewriting (rewrite_query) Repeating the earlier topic when you ask a follow-up "what if it fails?" after a VPN question: printer guide → VPN setup and ERR-4012
HyDE (hyde_search) Telling the librarian "I'm looking for a page that says something like…" "a weird message asking for my bank details": nothing relevant → phishing guide first
Multi-query (multi_query_search) Asking three librarians in three different phrasings "two step login setup": no MFA guide → MFA guide in top 3
Metadata filter (updated_after) Ignoring the out-of-date binder on the shelf superseded 2023 travel policy first → gone
Contextual retrieval (RAGIndex(contextual=True)) Writing the chapter title on every index card "how do I fix ERR-4012": fix sentence missing → found
Parent-child (retrieve_parents) Find the sentence, hand over the whole page a matching sentence → its full document

HyDE (Hypothetical Document Embeddings) deserves a picture, because it sounds backwards: you ask a model to invent an answer, then search with it.

flowchart LR Q["Question:<br/>'weird message asking<br/>for my bank details'"] --> M[Model drafts a<br/>plausible answer] M --> H["Draft: 'This sounds like a phishing<br/>email… report it in Outlook'"] H --> E[Embed the draft] E --> S[Dense search] S --> R[Phishing guide, rank 1]

Reading it: questions and answers are written differently. The user's words ("weird", "bank details") appear nowhere in the phishing guide, but the drafted answer uses the vocabulary answers use ("phishing", "report", "suspicious"), so its embedding lands next to the real passage. The draft may be wrong in its details; that's fine, because it's only used to search and is never shown to the user.

Contextual retrieval fixes the "second index card" problem from the chunking section: before indexing, each passage gets a short line saying where it comes from. Here that's just the title and department; Anthropic's version asks a model to write one sentence situating the chunk in its document.

Without a context line the fix passage misses the top 20 for all three phrasings; with the title prefixed it ranks 2nd, 2nd and 4th

Reading it: three ways of asking how to fix ERR-4012; each pair of bars is the rank of the passage "Update AnyConnect to version 5.1 or later and reboot." (shorter is better, and 21 means not in the top 20). Without context the passage never mentions ERR-4012, and for all three phrasings it doesn't make the top 20 at all. With the title prefixed it ranks 2nd, 2nd and 4th: in the top five every time, though "resolve ERR-4012 on my laptop" still ranks three passages from the laptop guide above it. A context line turns "never found" into "found", and ranking the right passage first is then the reranker's job.

Everyday picture. A researcher rather than a librarian: reads the question, decides what to look up, reads the results, and searches again if something's missing.

Plain RAG searches exactly once, with the user's words. In agentic RAG search is a tool the model calls: zero times for small talk, several times for a multi-part question, with its own reformulated queries.

Worked example. "What is the per-diem for meals and how do I upload expense receipts?" produces two searches ("What is the per-diem for meals", "how do I upload expense receipts") and an answer citing both fin-002#2 and fin-004#0. "thanks, that's all!" produces no search at all.

sequenceDiagram participant U as User participant M as Model participant S as search_kb tool U->>M: per-diem for meals AND how to upload receipts? M->>S: search_kb("per-diem for meals") S-->>M: fin-002 passages M->>S: search_kb("how do I upload expense receipts") S-->>M: fin-004 passages M-->>U: both answers, each with its citation

Reading it: the model, not a fixed pipeline, decides how many searches to run and with which words. That handles multi-part and vague questions better, at the cost of more model calls, more latency, and a loop that needs the usual step budget (primer.agents.agent_loop). The search tool here also applies the freshness filter itself, so the agent can't forget it.

In code: agentic_answer offers the model one search tool and runs the loop: each search retrieves, reranks and returns tagged passages, until the model answers or the step limit is hit.

GraphRAG: questions about connections

Everyday picture. A detective's corkboard: photos of people and places joined by string, each string labelled with the note that established it.

Some questions aren't answered by any single passage: "what depends on two-factor authentication?" or "what are the main themes across these contracts?". GraphRAG first extracts entities (named things: systems, people, products) and the relationships between them into a graph, then answers by walking it, or by summarizing groups of connected entities. build_graph does the smallest honest version: two entities are linked when one sentence mentions both, and each link remembers which document said so.

flowchart LR TF((two-factor)) ---|it-006| VPN((VPN)) TF ---|it-006| EM((email)) TF ---|it-006| AA((authenticator app)) TF ---|it-003| AC((AnyConnect)) TF ---|it-003| SC((software center)) TF ---|hr-003| LT((laptop)) AC ---|it-003| SC VPN ---|it-006| EM AA ---|it-001| SSP((self-service portal))

Reading it: each circle is an entity found in the knowledge base; each line is a sentence that mentioned both ends, labelled with its source document. "What depends on two-factor?" is answered by reading the neighbours of one circle, with citations for free. Real GraphRAG uses a model to extract typed relations ("requires", "replaces") and to write a summary for each cluster, which is what makes broad, whole-collection questions answerable.

In code: communities groups the graph from build_graph into connected clusters of entities, the "themes" a real GraphRAG system would summarize.

Debugging a confident wrong answer

Everyday picture. A student gets an exam question wrong. Either the librarian brought the wrong book, or the student misread the right one. The fixes are completely different, so find out which before changing anything.

Worked example. debug_wrong_answers runs the 12 labelled questions and classifies each: retrieval miss (no relevant document in the top 3: no prompt can fix it), generation miss (the right document was there but the answer didn't use it), or ok. With dense-only retrieval, "what does ERR-4012 mean" is a retrieval miss and recall@3 is 0.92; switching to hybrid fixes it (recall@3 = 1.00). What's left are generation misses. One of them, "per-diem for meals when traveling", is answered from the superseded 2023 policy: a data problem that a date filter fixes. The other, "enroll in MFA", is a limit of this lesson's stand-in model rather than of RAG: it matches words literally, so "enroll" misses the passage's "enrolling", and with only "MFA" in common it declines to answer. A real model would read it-006 and answer.

flowchart TD W[Confident wrong answer] --> R{Right passage in<br/>the top k?} R -->|no| P{Is it in the<br/>index at all?} P -->|no| FIXI[Fix ingestion:<br/>parsing, chunking, permissions sync] P -->|yes| FIXR[Fix retrieval:<br/>hybrid, rerank, rewrite, HyDE, filters] R -->|yes| C{Was it in the<br/>final context?} C -->|no| FIXC[Fix assembly:<br/>reranker, budget, top_n] C -->|yes| FIXG[Fix generation:<br/>instructions, citations, stale sources]

Reading it: start at the top and answer each question with a measurement, not a guess. Most teams jump straight to the bottom-right box and edit the prompt; the diagram says to rule out the three earlier failure points first, because they're more common and a prompt can't fix them.

Dense search has the only retrieval miss; switching to hybrid removes it, and the wrong answers left are generation misses

Reading it: each bar is one retrieval method over the 12 labelled questions, split into ok (green), generation misses (orange) and retrieval misses (red). Dense search has a retrieval miss (the error code); hybrid removes it. The orange that remains is where to look next, and it's not retrieval.

In 20 seconds

  • RAG retrieves relevant passages at question time and has the model answer from them with citations: knowledge that changes or must be cited, with no retraining.
  • Quality is decided early: parsing (tables!), chunking with metadata, and hybrid retrieval with reranking.
  • Filter by permissions before retrieval, never after generation.
  • Upgrades: query rewriting, HyDE, multi-query, date filters, contextual retrieval, parent-child; agentic RAG for multi-part questions; GraphRAG for questions about connections.
  • Debug by measuring retrieval (recall@k) first, then generation.

Self-test questions

Your RAG system gives confident wrong answers. Debug it step by step. Take the failing questions (and build a labelled set if you don't have one). First, is the right passage in the index at all? If not, it's ingestion: parsing, chunking or permission sync. Second, is it in the top k? Measure recall@k; if not, fix retrieval: hybrid search, a reranker, query rewriting or HyDE, metadata filters, contextual chunks, domain-tuned embeddings. Third, did it survive into the final context? If not, fix reranking or the context budget. Only then look at generation: instructions to answer only from sources and to decline otherwise, citation checks, stale or contradictory sources. Add every case to the golden set.

Why does hybrid search beat pure vector search on enterprise data? Enterprise questions are full of exact identifiers (error codes, product names, ticket numbers) that embeddings blur, and full of paraphrases that keyword search misses. Hybrid gets both, and RRF merges the rankings without having to reconcile incompatible score scales.

How do you make RAG respect document permissions? Copy each document's access-control list onto every chunk at ingestion, filter candidates by the user's groups before ranking, and sync permission changes from the source system promptly. Never filter after generation: the model has already used the restricted text.

When would you use agentic RAG instead of a single retrieval step? For multi-part or vague questions, and when the model needs to judge whether results are good enough and search again. It costs more calls and latency, so keep single-shot retrieval for simple lookups.

RAG or fine-tuning for a company knowledge assistant? RAG for knowledge: it updates by re-indexing, cites sources, and respects permissions. Fine-tuning teaches behaviour and format, not facts that change. Most systems are RAG plus a good prompt (see primer.ml.training_stages).

The papers behind this lesson

  • Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020), https://arxiv.org/abs/2005.11401. Coined the term and showed that pairing a generator with a retriever over a document index beats a model relying on its weights alone for knowledge-heavy questions. annotated companion
  • Gao, Ma, Lin & Callan, Precise Zero-Shot Dense Retrieval without Relevance Labels (2022), https://arxiv.org/abs/2212.10496. Introduced HyDE: search with an embedding of a model-written hypothetical answer. annotated companion
  • Edge et al., From Local to Global: A Graph RAG Approach to Query-Focused Summarization (2024), https://arxiv.org/abs/2404.16130. Builds an entity graph and community summaries so that whole-collection questions can be answered.

Further reading

on GitHub
   1r"""
   2# Retrieval-augmented generation (RAG), end to end
   3
   4Run: `python -m primer.agents.rag`
   5
   6This lesson builds on keyword search, embeddings and reranking from
   7`primer.ml.embeddings.retrieval`, on prompt assembly from
   8`primer.agents.context`, and on the tool-calling loop from
   9`primer.agents.agent_loop`.
  10
  11## Level 1: The practitioner's guide
  12
  13**In one sentence.** Retrieval-augmented generation (RAG) fetches the few
  14passages of your own documents most likely to answer a question, puts them
  15in the prompt, and has the model answer from those passages with citations,
  16so the model can use knowledge it was never trained on and a reader can
  17check where each fact came from.
  18
  19**When you need it.** When the answer lives in documents the model has not
  20seen (your policies, tickets, contracts, product manuals), when those
  21documents change faster than you could retrain, when every answer must name
  22its source, or when different users may read different documents. You don't
  23need it when the model already knows the subject (public, stable knowledge),
  24when you want to change *how* the model behaves rather than *what* it
  25knows (that is fine-tuning: it teaches format and style, not facts that
  26change), or when the whole collection fits in the prompt. Anthropic's
  27contextual retrieval post puts that last threshold at about 200,000 tokens,
  28roughly 500 pages: below it, put everything in the prompt and skip the
  29pipeline. The tell: your model answers a question about the 2023 travel
  30policy confidently and wrongly, because it never saw the 2026 one.
  31
  32**Your options.** From the cheapest to the most capable, each one usually
  33added on top of the last:
  34
  35| Option | What it does | What it guarantees | What it costs | Where it lives |
  36|---|---|---|---|---|
  37| Everything in the prompt | Sends the whole collection with every question | Nothing is ever missed by a search | Every token, every call; only up to a few hundred pages | Your prompt |
  38| Keyword search (BM25) | Ranks passages by the question's rarer words | Exact identifiers (`ERR-4012`, ticket numbers) are found | An inverted index; misses synonyms and paraphrase | A search engine |
  39| Dense search | Ranks passages by embedding similarity, meaning over words | Paraphrases are found ("scam message" finds the phishing guide) | An embedding model at ingest and per query, a vector index; blurs exact codes | A vector index |
  40| Hybrid search | Runs both and fuses the rankings by rank (RRF), never by score | Both kinds of question, and no score scales to reconcile | Two indexes, two searches per question | Both indexes plus a fusion step |
  41| Reranking | A cross-encoder reads the question next to each of the top candidates | Far better ordering of the top few | One model score per candidate, so only over a shortlist | A reranking model between search and the prompt |
  42| Retrieval upgrades | Contextual chunks, query rewriting, HyDE, multi-query, date filters, parent-child | Each fixes one named failure | Model calls at ingest (context lines) or per query (rewrites, drafts) | Ingestion or query time |
  43| Agentic RAG | The model decides when, how often and with what words to search | Multi-part and vague questions handled; no search for small talk | Several model calls and a step budget per question | An agent loop with a search tool |
  44| GraphRAG | Extracts entities and relations first, answers by walking the graph or summarizing clusters | Questions about connections and whole-collection themes | A model pass over every document at ingest, and a graph to maintain | Ingestion, plus a graph store |
  45
  46**How to choose.** Start from the questions people actually ask, with a
  47labelled set of at least a dozen where you know the right document.
  48
  49- Questions full of exact identifiers and paraphrases both, which is what
  50  enterprise data looks like: hybrid search with a reranker. In this
  51  lesson's toy set, dense search alone never finds `ERR-4012` and plateaus
  52  at 92% recall; hybrid reaches 100% by the second result; reranking the
  53  hybrid shortlist lifts the share of questions answered by the first
  54  result from 75% to 92%.
  55- Chunks that lose their meaning when cut from the document (a fix that
  56  never names the error it fixes): contextual chunks. Prefixing each chunk
  57  with where it comes from turns "not in the top 20" into "second" here;
  58  Anthropic measured a 49% drop in top-20 retrieval failures from
  59  contextual embeddings and contextual BM25, and 67% with a reranker added.
  60- Follow-up questions and vague phrasing: query rewriting and HyDE, which
  61  searches with a model-written draft answer.
  62- Two questions in one, or a question the first search doesn't settle:
  63  agentic RAG.
  64- "What depends on X?" and "what are the themes?": GraphRAG, and only then,
  65  because it is the most expensive to build and keep current.
  66- Whatever you pick: filter by permissions and date *before* ranking, and
  67  verify citations before the answer reaches anyone. Retrieval quality is
  68  decided in the first boxes (parsing, chunking, search), never by the
  69  prompt.
  70
  71**What it costs.** Ingestion costs an embedding per chunk, a keyword index,
  72and for contextual chunks a model call per chunk: with prompt caching,
  73Anthropic's post prices that at \$1.02 per million document tokens. A query
  74costs one embedding, two searches, a reranker score for each of a handful
  75of candidates (8 here, kept to 3), and one generation whose input is the
  76question plus those passages; that prompt is what the answer's latency and
  77token bill mostly consist of. Agentic RAG multiplies the model calls by the
  78number of searches (two for the two-part question in this lesson, zero for
  79"thanks, that's all!") and needs a step limit. Quality costs come from the
  80front of the pipeline: a table flattened by a naive PDF extractor, so that
  81the hotel cap floats free of its grade, is a failure no later stage can
  82repair. Effort goes mostly into the labelled question set and the parsing,
  83not the model.
  84
  85**What breaks.**
  86
  87- **The right passage is not in the index.** Parsing garbled it, chunking
  88  split it from its context, or the permission sync dropped it. Check
  89  ingestion before touching anything downstream.
  90- **The right passage is in the index but not in the top k.** Measure
  91  recall@k on the labelled set. If it is low, change retrieval (hybrid,
  92  reranker, rewriting, HyDE, contextual chunks); a prompt cannot fix it.
  93- **Permission leaks.** Filtering after generation strips the citation and
  94  leaves the fact: in this lesson the restricted forecast's "41 million
  95  dollars" reaches an employee who may not read it. Copy each document's
  96  access list onto its chunks and filter before ranking; sync revocations
  97  promptly.
  98- **Stale sources.** A superseded 2023 policy outranks the current one.
  99  Carry the date on every chunk and filter or prefer by it.
 100- **Confident answers with no support.** Instruct the model to answer only
 101  from the sources and to decline otherwise, then check that every cited id
 102  was in the context and that its passage supports the claim.
 103- **Blaming the model.** With dense-only retrieval this lesson's twelve
 104  questions produce five wrong answers, of which only one is a retrieval
 105  miss; switching to hybrid removes that one, and the two that remain are
 106  generation and data problems. Each needs a different fix, so diagnose
 107  first.
 108
 109**In the wild.** The name comes from Lewis et al. (2020), who paired a
 110retriever with a generator over a document index. Keyword search is BM25
 111in Lucene, Elasticsearch and OpenSearch; dense search runs on vector
 112indexes such as FAISS and hosted services such as Pinecone; Elasticsearch
 113documents RRF for fusing the two. HyDE is Gao et al. (2022). Contextual
 114retrieval and prompt caching are described in Anthropic's post, and
 115Claude's citations feature returns the exact passage behind each claim for
 116PDF, plain-text and custom documents. GraphRAG is Microsoft's open
 117implementation of Edge et al. (2024). RAGAS evaluates a RAG pipeline on
 118retrieval and generation separately, which is exactly the split the
 119debugging section relies on.
 120
 121**Go deeper.** Level 2 builds the whole pipeline in plain Python: a PDF
 122extractor's damage repaired step by step, chunks that carry their
 123metadata, BM25 and dense search fused by a one-line formula, a reranker, a
 124prompt that tags every source, a citation check, the permission leak
 125reproduced, each upgrade as a working function with its before and after,
 126and the diagnosis figure you can rerun on your own questions. If you only
 127needed to choose, you are done.
 128
 129## Level 2: How it works, from scratch
 130
 131What follows builds every box of the pipeline from scratch, in the order a
 132question travels through them, then the upgrades, then how to tell a
 133retrieval failure from a generation failure.
 134
 135An open-book exam with a librarian. Before you answer, the librarian fetches
 136the few pages most likely to contain the answer; you answer from those
 137pages and write down which page each fact came from. If the pages don't
 138contain the answer, you say so instead of guessing.
 139
 140**Retrieval-augmented generation** is that arrangement for a language
 141model. The model's training knowledge is frozen and doesn't include your
 142company's documents, so before each answer the system *retrieves* relevant
 143passages from your documents and puts them in the prompt, and the model
 144*generates* an answer from them, with citations. It's how you give a model
 145knowledge that changes, or that must be cited, without retraining it.
 146
 147```mermaid
 148flowchart LR
 149  subgraph ING["Ingestion, ahead of time"]
 150    D[Documents] --> P[Parse] --> C[Chunk + metadata] --> E[Embed + keyword index]
 151  end
 152  E --> IX[(Index)]
 153  subgraph QRY["Query time, per question"]
 154    Q[Question] --> F[Filter by permissions<br/>and date] --> R[Retrieve: hybrid] --> RR[Rerank] --> AC[Assemble context] --> G[Generate + cite] --> V[Verify citations]
 155  end
 156  IX --> F
 157```
 158
 159**Reading it:** the top row runs once per document, whenever documents
 160change; the bottom row runs for every question. Everything the answer can
 161contain has to survive every box, left to right: a table mangled at *Parse*
 162or a passage missing at *Retrieve* can't be fixed by a better prompt at
 163*Generate*. That's why most RAG quality work happens in the first boxes, not
 164the last.
 165
 166**In code:** `RAGIndex` is the top row: it chunks every document and builds
 167the keyword and dense indexes. `answer` is the bottom row, from filtered
 168retrieval to verified citations, and returns a `RAGResult`.
 169
 170The rest of this lesson walks the boxes in order, then covers the upgrades
 171and how to debug a wrong answer.
 172
 173## Parsing: where quality dies first
 174
 175**Everyday picture.** Photocopy a newspaper page and read it straight
 176across: you get the first line of one column, then the first line of the
 177next story, and nothing makes sense. Text extracted from PDFs has the same
 178problems: running headers and page numbers mixed into the text, words
 179broken across lines, and tables flattened into a row of numbers with no
 180columns.
 181
 182**Worked example.** `MESSY_PDF` is two pages of a travel policy as a naive
 183extractor returns it. `clean_extracted_text` repairs it in four steps:
 184
 185| Damage | Before | After |
 186|---|---|---|
 187| Header repeated on every page | `Northwind Ltd  Travel Policy 2026 … CONFIDENTIAL` | removed (it appears on 2 of 2 pages) |
 188| Page-number footer | `Page 1 of 2` | removed |
 189| Word hyphenated across a line break | `all em-` / `ployees travelling` | `all employees travelling` |
 190| Table flattened | `L1-L3 150 60 L4-L6 200 75` | kept as rows, then one sentence per row |
 191
 192The table is the dangerous one. Flattened, "200" floats free of its grade.
 193`table_rows_as_sentences` turns each row into a self-contained sentence,
 194`Grade L4-L6: hotel cap 200, meal cap 75.`, so any row can be retrieved on
 195its own and still says what its numbers mean.
 196
 197```mermaid
 198flowchart LR
 199  RAW[Raw extracted text] --> H[Drop lines repeated<br/>on most pages]
 200  H --> N[Drop page numbers]
 201  N --> J[Rejoin hyphenated words,<br/>unwrap paragraphs]
 202  J --> T{Table rows?}
 203  T -->|yes| S[One sentence per row,<br/>headers repeated]
 204  T -->|no| P[Paragraphs]
 205  S --> OUT[Clean text for chunking]
 206  P --> OUT
 207```
 208
 209**Reading it:** each box removes one kind of extraction damage. The diamond
 210is the important decision: prose and tables need different handling, because
 211a table's meaning lives in its columns. Real pipelines add layout-aware
 212parsers and **OCR** (optical character recognition: reading text from an
 213image of a page) for scanned documents, with a quality check on the output.
 214
 215**Why it matters.** Enterprise documents are scanned, multi-column, full of
 216tables that span pages. If parsing garbles the numbers, no retriever or
 217model can recover them, and the failure looks like a *model* error.
 218
 219**In code:** `extract_table` finds a table by its header row in the cleaned
 220text and returns each row as a dict, ready for `table_rows_as_sentences`.
 221
 222## Chunking, with metadata that travels
 223
 224**Everyday picture.** Index cards. Each card holds one idea, small enough to
 225match a question precisely, and in its corner a label: which document,
 226which date, who may read it.
 227
 228**Worked example.** With a 15-word limit, `it-004` splits into two passages:
 229
 230| Passage id | Text | Updated | Readers |
 231|---|---|---|---|
 232| `it-004#0` | ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. | 2026-04-10 | everyone |
 233| `it-004#1` | Update AnyConnect to version 5.1 or later and reboot. | 2026-04-10 | everyone |
 234
 235(The limit is tiny because the toy documents are tiny; real systems use a few
 236hundred tokens.) Split on the document's structure (sentences, paragraphs,
 237sections), not at fixed character counts that cut a sentence in half.
 238`primer.ml.embeddings.retrieval` compares chunking strategies in depth.
 239
 240**Why it matters.** Small chunks match precisely but can lose context (the
 241second card doesn't say *which* error it fixes; see contextual retrieval
 242below). The metadata is what makes permission and date filters possible
 243later.
 244
 245**In code:** `chunk_document` splits a document on sentence boundaries into
 246`Passage`s, and each `Passage` carries its id, title, department, date and
 247readers.
 248
 249## Retrieval: keyword and meaning, fused
 250
 251**Everyday picture.** Two librarians: one searches the catalogue for your
 252exact words, the other understands what you *mean* even when you use
 253different words. You take the books both of them rank highly.
 254
 255* **Keyword search (BM25)** scores passages by the question's words,
 256  giving more weight to rare words and less to long passages. It nails
 257  exact identifiers like `ERR-4012` and misses synonyms.
 258* **Dense search** compares **embeddings** (vectors where closeness means
 259  similar meaning; see `primer.ml.embeddings`). It finds "scam message in my
 260  inbox" → the phishing guide with no shared words, but blurs exact codes.
 261* **Hybrid search** runs both and merges the two rankings with **reciprocal
 262  rank fusion (RRF)**, which only needs the ranks, never the scores. BM25
 263  scores and cosine similarities live on different scales, so adding them
 264  would be meaningless.
 265
 266$$
 267\text{RRF}(d) = \sum_{i} \frac{1}{k + \text{rank}_i(d)}
 268$$
 269
 270**Symbols**
 271
 272| Symbol | Meaning here | Worked example |
 273|---|---|---|
 274| $d$ | one passage | it-004#0 |
 275| $i$ | one of the rankings being fused (keyword, dense) | 2 rankings |
 276| $\text{rank}_i(d)$ | position of $d$ in ranking $i$ (1 = best); rankings that miss $d$ contribute nothing | 1st, 3rd |
 277| $k$ | a damping constant; 60 is the usual choice | 60 |
 278| $\sum_i$ | add up over the rankings | |
 279
 280**In words:** each ranking gives a passage a vote worth one over sixty-plus
 281its position, and the votes are added.
 282
 283**On the worked example:** a passage ranked 1st by dense search and 3rd by
 284keyword search scores 1/61 + 1/63 = 0.0323. A passage ranked 1st by one
 285list and absent from the other scores only 1/61 = 0.0164. Passages that both
 286methods agree on rise to the top.
 287
 288**In Python:**
 289
 290```python
 291k = 60
 292# 1st by dense, 3rd by keyword: Σ_i 1/(k + rank_i(d))
 293round(1 / (k + 1) + 1 / (k + 3), 4)  # → 0.0323
 294# 1st in one list, absent from the other
 295round(1 / (k + 1), 4)  # → 0.0164
 296```
 297
 298![Dense search plateaus at 92% recall because it never finds the error code, while hybrid reaches 100% by k = 2 and reranking lifts its top-1 from 75% to 92%](figures/primer.agents.rag.recall.svg)
 299
 300**Reading it:** the x-axis is how many passages you keep (k); the y-axis is
 301**recall@k**, the share of the 12 labelled questions whose correct document
 302is among those k. On this small set many questions reuse the documents'
 303own words, so keyword search is already strong. Dense search plateaus at
 30492% because it never finds `ERR-4012`. Hybrid reaches 100% by k = 2 but only
 30575% at k = 1. Reranking the hybrid shortlist lifts the top-1 result to 92%.
 306Run this measurement on *your* questions before choosing a method.
 307
 308$$
 309\text{recall@}k = \frac{1}{|Q|} \sum_{q \in Q} \mathbb{1}\bigl[\text{a relevant document is in the top } k \text{ for } q\bigr]
 310$$
 311
 312**Symbols**
 313
 314| Symbol | Meaning here |
 315|---|---|
 316| $Q$ | the labelled questions; $\lvert Q\rvert$ is how many (12 here) |
 317| $q$ | one question |
 318| $\mathbb{1}[\ldots]$ | 1 if true, 0 if not |
 319
 320**In words:** the share of questions for which the right document made it
 321into the top k.
 322
 323**On the worked example:** dense search at k = 3 finds the right document for
 32411 of 12 questions (all but `ERR-4012`): 11/12 = 0.92.
 325
 326**In Python:**
 327
 328```python
 329# 𝟙[...] for each of the 12 questions; ERR-4012 is the 0
 330found = [1] * 11 + [0]
 331# (1/|Q|) Σ over q in Q
 332round(sum(found) / len(found), 2)  # → 0.92
 333```
 334
 335**In code:** `RAGIndex.retrieve` ranks the allowed passages with
 336`primer.ml.embeddings.retrieval.BM25`, with dense similarity, or with both
 337fused by `primer.ml.embeddings.retrieval.reciprocal_rank_fusion`, depending
 338on the mode you ask for. `recall_curve` measures recall@k for each method.
 339
 340## Reranking: read the question and each passage together
 341
 342**Everyday picture.** The librarians bring back twenty books quickly; an
 343expert then reads your question next to each book and picks the best three.
 344
 345A **reranker** (usually a **cross-encoder**: a model that reads the question
 346and one passage *together* and outputs a relevance score) is far more
 347precise than comparing precomputed vectors, and far too slow to run over
 348every passage. So retrieval casts a wide, cheap net (here 8 candidates) and
 349the reranker keeps the best few (here 3). This module reuses the toy
 350cross-encoder from `primer.ml.embeddings.retrieval`.
 351
 352**In code:** `RAGIndex.rerank` scores each (question, passage) pair with
 353`primer.ml.embeddings.retrieval.CrossEncoder` and keeps the best few.
 354
 355## Assemble, generate, cite, verify
 356
 357**Worked example.** "What does ERR-4012 mean?" becomes this context (each
 358source escaped and tagged with its id; the question last; see
 359`primer.agents.context`):
 360
 361```text
 362<sources>
 363<source id="it-004#1" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">Update AnyConnect to version 5.1 or later and reboot.</source>
 364<source id="it-004#0" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date.</source>
 365<source id="it-002#2" title="Password policy" updated="2025-11-20">This policy applies to all employees and contractors.</source>
 366</sources>
 367<question>What does ERR-4012 mean?</question>
 368```
 369
 370and the answer is *"ERR-4012 means the VPN tunnel could not be established,
 371usually because the client is out of date. [it-004#0]"*. `verify_citations`
 372then checks that every cited id was really in the context and that the
 373cited passage supports the claim before it; a question the sources don't
 374cover gets "I don't know based on the provided sources."
 375
 376```mermaid
 377sequenceDiagram
 378  participant U as User
 379  participant A as RAG service
 380  participant I as Index
 381  participant M as Model
 382  U->>A: What does ERR-4012 mean?
 383  A->>I: hybrid search, filtered to what this user may read
 384  I-->>A: 8 candidate passages
 385  A->>A: rerank, keep 3
 386  A->>M: rules + tagged sources + question
 387  M-->>A: answer with [it-004#0]
 388  A->>A: verify: id was in context, passage supports claim
 389  A-->>U: answer + clickable citation
 390```
 391
 392**Reading it:** time runs downward. Note the order: the permission filter
 393runs at the index, before any passage is fetched; the model only ever sees
 394three passages; and nothing reaches the user until the citations check out.
 395Citations are what let a person verify an answer in seconds, which is most
 396of what makes people trust the system.
 397
 398**In code:** `assemble_context` builds the tagged sources and the question.
 399`generate` sends them to the model with the answer-from-sources rules, and
 400`grounded_policy` is the offline stand-in model that honours them. `answer`
 401chains retrieve, rerank, assemble, generate and `verify_citations`.
 402
 403## Permission-aware retrieval: filter before, never after
 404
 405**Everyday picture.** A librarian checks your library card *before*
 406fetching from the restricted archive. Checking it after you've read the
 407document is pointless.
 408
 409Each passage carries its **access-control list** (ACL: the groups allowed to
 410read it), copied from the source system (SharePoint, Google Drive, a wiki).
 411`RAGIndex.retrieve` drops every passage the user can't read *before*
 412ranking, so restricted text can't reach the context, the model, or the
 413answer.
 414
 415**Worked example.** "What is the Q3 revenue forecast?" (the forecast
 416document is readable only by `finance` and `exec`):
 417
 418| Who asks | How it's filtered | Answer |
 419|---|---|---|
 420| finance user | before retrieval | "Q3 revenue is forecast at 41 million dollars… [fin-006#0]" |
 421| anyone else | before retrieval | "I don't know based on the provided sources." |
 422| anyone else | **after** generation (citation removed) | "Q3 revenue is forecast at **41 million dollars**…" |
 423
 424```mermaid
 425sequenceDiagram
 426  participant U as Employee (not finance)
 427  participant A as RAG service
 428  participant M as Model
 429  rect rgb(250, 225, 225)
 430  Note over A,M: Wrong: filter after generation
 431  A->>M: sources include the restricted forecast
 432  M-->>A: "…41 million dollars [fin-006#0]"
 433  A->>A: strip the restricted citation
 434  A-->>U: "…41 million dollars" (leaked)
 435  end
 436  rect rgb(225, 240, 225)
 437  Note over A,M: Right: filter before retrieval
 438  A->>A: drop passages the user can't read
 439  A->>M: sources without the forecast
 440  M-->>A: "I don't know…"
 441  A-->>U: nothing restricted
 442  end
 443```
 444
 445**Reading it:** in the red box the model saw the restricted passage, so the
 446number is already in its words, and deleting the citation marker afterwards
 447leaves the fact behind. In the green box the passage never left the index.
 448Permission changes in the source system must also sync to the index
 449quickly, or a revoked user keeps access until the next re-index.
 450
 451**In code:** `post_generation_filter_answer` is the wrong way, kept so the
 452leak can be shown: it retrieves for every group, generates, then strips the
 453restricted citation markers.
 454
 455## Retrieval upgrades
 456
 457Each upgrade below is a working function, and each fixes a failure you can
 458reproduce in `demo()`:
 459
 460| Upgrade | Everyday picture | Before → after (top 3) |
 461|---|---|---|
 462| **Query rewriting** (`rewrite_query`) | Repeating the earlier topic when you ask a follow-up | "what if it fails?" after a VPN question: printer guide → VPN setup and ERR-4012 |
 463| **HyDE** (`hyde_search`) | Telling the librarian "I'm looking for a page that says something like…" | "a weird message asking for my bank details": nothing relevant → phishing guide first |
 464| **Multi-query** (`multi_query_search`) | Asking three librarians in three different phrasings | "two step login setup": no MFA guide → MFA guide in top 3 |
 465| **Metadata filter** (`updated_after`) | Ignoring the out-of-date binder on the shelf | superseded 2023 travel policy first → gone |
 466| **Contextual retrieval** (`RAGIndex(contextual=True)`) | Writing the chapter title on every index card | "how do I fix ERR-4012": fix sentence missing → found |
 467| **Parent-child** (`retrieve_parents`) | Find the sentence, hand over the whole page | a matching sentence → its full document |
 468
 469**HyDE** (Hypothetical Document Embeddings) deserves a picture, because it
 470sounds backwards: you ask a model to *invent* an answer, then search with
 471it.
 472
 473```mermaid
 474flowchart LR
 475  Q["Question:<br/>'weird message asking<br/>for my bank details'"] --> M[Model drafts a<br/>plausible answer]
 476  M --> H["Draft: 'This sounds like a phishing<br/>email… report it in Outlook'"]
 477  H --> E[Embed the draft]
 478  E --> S[Dense search]
 479  S --> R[Phishing guide, rank 1]
 480```
 481
 482**Reading it:** questions and answers are written differently. The user's
 483words ("weird", "bank details") appear nowhere in the phishing guide, but
 484the drafted answer uses the vocabulary answers use ("phishing", "report",
 485"suspicious"), so its embedding lands next to the real passage. The draft
 486may be wrong in its details; that's fine, because it's only used to search
 487and is never shown to the user.
 488
 489**Contextual retrieval** fixes the "second index card" problem from the
 490chunking section: before indexing, each passage gets a short line saying
 491where it comes from. Here that's just the title and department; Anthropic's
 492version asks a model to write one sentence situating the chunk in its
 493document.
 494
 495![Without a context line the fix passage misses the top 20 for all three phrasings; with the title prefixed it ranks 2nd, 2nd and 4th](figures/primer.agents.rag.contextual.svg)
 496
 497**Reading it:** three ways of asking how to fix ERR-4012; each pair of bars
 498is the rank of the passage "Update AnyConnect to version 5.1 or later and
 499reboot." (shorter is better, and 21 means not in the top 20). Without
 500context the passage never mentions ERR-4012, and for all three phrasings it
 501doesn't make the top 20 at all. With the title prefixed it ranks 2nd, 2nd
 502and 4th: in the top five every time, though "resolve ERR-4012 on my
 503laptop" still ranks three passages from the laptop guide above it. A context line turns
 504"never found" into "found", and ranking the right passage first is then
 505the reranker's job.
 506
 507## Agentic RAG: let the model decide when and what to search
 508
 509**Everyday picture.** A researcher rather than a librarian: reads the
 510question, decides what to look up, reads the results, and searches again if
 511something's missing.
 512
 513Plain RAG searches exactly once, with the user's words. In **agentic RAG**
 514search is a tool the model calls: zero times for small talk, several times
 515for a multi-part question, with its own reformulated queries.
 516
 517**Worked example.** "What is the per-diem for meals and how do I upload
 518expense receipts?" produces two searches ("What is the per-diem for meals",
 519"how do I upload expense receipts") and an answer citing both `fin-002#2`
 520and `fin-004#0`. "thanks, that's all!" produces no search at all.
 521
 522```mermaid
 523sequenceDiagram
 524  participant U as User
 525  participant M as Model
 526  participant S as search_kb tool
 527  U->>M: per-diem for meals AND how to upload receipts?
 528  M->>S: search_kb("per-diem for meals")
 529  S-->>M: fin-002 passages
 530  M->>S: search_kb("how do I upload expense receipts")
 531  S-->>M: fin-004 passages
 532  M-->>U: both answers, each with its citation
 533```
 534
 535**Reading it:** the model, not a fixed pipeline, decides how many searches
 536to run and with which words. That handles multi-part and vague questions
 537better, at the cost of more model calls, more latency, and a loop that needs
 538the usual step budget (`primer.agents.agent_loop`). The search tool here also
 539applies the freshness filter itself, so the agent can't forget it.
 540
 541**In code:** `agentic_answer` offers the model one search tool and runs the
 542loop: each search retrieves, reranks and returns tagged passages, until the
 543model answers or the step limit is hit.
 544
 545## GraphRAG: questions about connections
 546
 547**Everyday picture.** A detective's corkboard: photos of people and places
 548joined by string, each string labelled with the note that established it.
 549
 550Some questions aren't answered by any single passage: "what depends on
 551two-factor authentication?" or "what are the main themes across these
 552contracts?". **GraphRAG** first extracts **entities** (named things:
 553systems, people, products) and the relationships between them into a graph,
 554then answers by walking it, or by summarizing groups of connected entities.
 555`build_graph` does the smallest honest version: two entities are linked when
 556one sentence mentions both, and each link remembers which document said so.
 557
 558```mermaid
 559flowchart LR
 560  TF((two-factor)) ---|it-006| VPN((VPN))
 561  TF ---|it-006| EM((email))
 562  TF ---|it-006| AA((authenticator app))
 563  TF ---|it-003| AC((AnyConnect))
 564  TF ---|it-003| SC((software center))
 565  TF ---|hr-003| LT((laptop))
 566  AC ---|it-003| SC
 567  VPN ---|it-006| EM
 568  AA ---|it-001| SSP((self-service portal))
 569```
 570
 571**Reading it:** each circle is an entity found in the knowledge base; each
 572line is a sentence that mentioned both ends, labelled with its source
 573document. "What depends on two-factor?" is answered by reading the
 574neighbours of one circle, with citations for free. Real GraphRAG uses a
 575model to extract typed relations ("requires", "replaces") and to write a
 576summary for each cluster, which is what makes broad, whole-collection
 577questions answerable.
 578
 579**In code:** `communities` groups the graph from `build_graph` into
 580connected clusters of entities, the "themes" a real GraphRAG system would
 581summarize.
 582
 583## Debugging a confident wrong answer
 584
 585**Everyday picture.** A student gets an exam question wrong. Either the
 586librarian brought the wrong book, or the student misread the right one. The
 587fixes are completely different, so find out which before changing anything.
 588
 589**Worked example.** `debug_wrong_answers` runs the 12 labelled questions and
 590classifies each: **retrieval miss** (no relevant document in the top 3: no
 591prompt can fix it), **generation miss** (the right document was there but the
 592answer didn't use it), or ok. With dense-only retrieval, "what does ERR-4012
 593mean" is a retrieval miss and recall@3 is 0.92; switching to hybrid fixes it
 594(recall@3 = 1.00). What's left are generation misses. One of them,
 595"per-diem for meals when traveling", is answered from the superseded 2023
 596policy: a *data* problem that a date filter fixes. The other, "enroll in
 597MFA", is a limit of this lesson's stand-in model rather than of RAG: it
 598matches words literally, so "enroll" misses the passage's "enrolling", and
 599with only "MFA" in common it declines to answer. A real model would read
 600it-006 and answer.
 601
 602```mermaid
 603flowchart TD
 604  W[Confident wrong answer] --> R{Right passage in<br/>the top k?}
 605  R -->|no| P{Is it in the<br/>index at all?}
 606  P -->|no| FIXI[Fix ingestion:<br/>parsing, chunking, permissions sync]
 607  P -->|yes| FIXR[Fix retrieval:<br/>hybrid, rerank, rewrite, HyDE, filters]
 608  R -->|yes| C{Was it in the<br/>final context?}
 609  C -->|no| FIXC[Fix assembly:<br/>reranker, budget, top_n]
 610  C -->|yes| FIXG[Fix generation:<br/>instructions, citations, stale sources]
 611```
 612
 613**Reading it:** start at the top and answer each question with a
 614measurement, not a guess. Most teams jump straight to the bottom-right box
 615and edit the prompt; the diagram says to rule out the three earlier failure
 616points first, because they're more common and a prompt can't fix them.
 617
 618![Dense search has the only retrieval miss; switching to hybrid removes it, and the wrong answers left are generation misses](figures/primer.agents.rag.diagnosis.svg)
 619
 620**Reading it:** each bar is one retrieval method over the 12 labelled
 621questions, split into ok (green), generation misses (orange) and retrieval
 622misses (red). Dense search has a retrieval miss (the error code); hybrid
 623removes it. The orange that remains is where to look next, and it's not
 624retrieval.
 625
 626## In 20 seconds
 627- RAG retrieves relevant passages at question time and has the model answer
 628  from them with citations: knowledge that changes or must be cited, with no
 629  retraining.
 630- Quality is decided early: parsing (tables!), chunking with metadata, and
 631  hybrid retrieval with reranking.
 632- Filter by permissions before retrieval, never after generation.
 633- Upgrades: query rewriting, HyDE, multi-query, date filters, contextual
 634  retrieval, parent-child; agentic RAG for multi-part questions; GraphRAG
 635  for questions about connections.
 636- Debug by measuring retrieval (recall@k) first, then generation.
 637
 638## Self-test questions
 639
 640**Your RAG system gives confident wrong answers. Debug it step by step.**
 641Take the failing questions (and build a labelled set if you don't have one).
 642First, is the right passage in the index at all? If not, it's ingestion:
 643parsing, chunking or permission sync. Second, is it in the top k? Measure
 644recall@k; if not, fix retrieval: hybrid search, a reranker, query rewriting
 645or HyDE, metadata filters, contextual chunks, domain-tuned embeddings.
 646Third, did it survive into the final context? If not, fix reranking or the
 647context budget. Only then look at generation: instructions to answer only
 648from sources and to decline otherwise, citation checks, stale or
 649contradictory sources. Add every case to the golden set.
 650
 651**Why does hybrid search beat pure vector search on enterprise data?**
 652Enterprise questions are full of exact identifiers (error codes, product
 653names, ticket numbers) that embeddings blur, and full of paraphrases that
 654keyword search misses. Hybrid gets both, and RRF merges the rankings without
 655having to reconcile incompatible score scales.
 656
 657**How do you make RAG respect document permissions?**
 658Copy each document's access-control list onto every chunk at ingestion,
 659filter candidates by the user's groups before ranking, and sync permission
 660changes from the source system promptly. Never filter after generation: the
 661model has already used the restricted text.
 662
 663**When would you use agentic RAG instead of a single retrieval step?**
 664For multi-part or vague questions, and when the model needs to judge whether
 665results are good enough and search again. It costs more calls and latency,
 666so keep single-shot retrieval for simple lookups.
 667
 668**RAG or fine-tuning for a company knowledge assistant?**
 669RAG for knowledge: it updates by re-indexing, cites sources, and respects
 670permissions. Fine-tuning teaches behaviour and format, not facts that
 671change. Most systems are RAG plus a good prompt (see
 672`primer.ml.training_stages`).
 673
 674## The papers behind this lesson
 675
 676- Lewis et al., *Retrieval-Augmented Generation for Knowledge-Intensive NLP
 677  Tasks* (2020), https://arxiv.org/abs/2005.11401. Coined the term and
 678  showed that pairing a generator with a retriever over a document index
 679  beats a model relying on its weights alone for knowledge-heavy questions.
 680  [annotated companion](../../papers/rag.html)
 681- Gao, Ma, Lin & Callan, *Precise Zero-Shot Dense Retrieval without
 682  Relevance Labels* (2022), https://arxiv.org/abs/2212.10496. Introduced
 683  HyDE: search with an embedding of a model-written hypothetical answer.
 684  [annotated companion](../../papers/hyde.html)
 685- Edge et al., *From Local to Global: A Graph RAG Approach to Query-Focused
 686  Summarization* (2024), https://arxiv.org/abs/2404.16130. Builds an entity
 687  graph and community summaries so that whole-collection questions can be
 688  answered.
 689
 690## Further reading
 691- Anthropic, *Introducing Contextual Retrieval*: https://www.anthropic.com/news/contextual-retrieval
 692- Anthropic citations docs: https://docs.claude.com/en/docs/build-with-claude/citations
 693- Microsoft GraphRAG documentation: https://microsoft.github.io/graphrag/
 694- RAGAS documentation (RAG evaluation): https://docs.ragas.io/
 695- The retrieval lesson in this repo, with BM25, RRF, cross-encoders and chunking in depth: `primer.ml.embeddings.retrieval`
 696"""
 697
 698from __future__ import annotations
 699
 700import re
 701from dataclasses import dataclass, field
 702from typing import Any
 703
 704import numpy as np
 705
 706from primer._show import banner, say, table, takeaway
 707from primer.agents.context import xml_wrap
 708from primer.agents.guardrails import claim_support
 709from primer.agents.llm import ScriptedLLM, ToolCall, last_user_text, tool_result_block, tool_results
 710from primer.common.corpus import DOCS, DOCS_BY_ID, LABELED_QUERIES, Doc
 711from primer.common.embedder import ConceptEmbedder
 712from primer.common.text import tokenize
 713from primer.ml.embeddings.retrieval import BM25, CrossEncoder, ideas, reciprocal_rank_fusion
 714
 715# ===========================================================================
 716# 1. INGESTION: parse, chunk, index
 717# ===========================================================================
 718
 719# What a PDF looks like after naive text extraction: a header and footer on
 720# every page, words hyphenated across line breaks, and a table whose columns
 721# survive only as runs of spaces.
 722MESSY_PDF = (
 723    "Northwind Ltd  Travel Policy 2026                    CONFIDENTIAL\n"
 724    "Hotel and meal limits apply to all em-\n"
 725    "ployees travelling on company busi-\n"
 726    "ness. Book through the travel portal.\n"
 727    "Grade    Hotel cap    Meal cap\n"
 728    "L1-L3    150          60\n"
 729    "L4-L6    200          75\n"
 730    "Page 1 of 2\n"
 731    "\f"
 732    "Northwind Ltd  Travel Policy 2026                    CONFIDENTIAL\n"
 733    "Economy class is required for flights un-\n"
 734    "der six hours.\n"
 735    "Page 2 of 2\n"
 736)
 737
 738_PAGE_NUMBER = re.compile(r"^page \d+ of \d+$", re.I)
 739_COLUMNS = re.compile(r"\s{2,}")  # two or more spaces separate table columns
 740
 741
 742def clean_extracted_text(raw: str) -> str:
 743    """Repair the usual damage of PDF text extraction.
 744
 745    1. Drop lines repeated on two or more pages (running headers/footers).
 746    2. Drop "Page n of m" footers.
 747    3. Rejoin words hyphenated across line breaks ("em-" + "ployees").
 748    4. Unwrap prose lines into paragraphs; keep table rows on their own lines.
 749    """
 750    pages = [p.strip("\n").split("\n") for p in raw.split("\f")]
 751    seen_on: dict[str, int] = {}
 752    for lines in pages:
 753        for line in set(l.strip() for l in lines):
 754            seen_on[line] = seen_on.get(line, 0) + 1
 755    kept = [l.strip() for lines in pages for l in lines
 756            if l.strip() and seen_on[l.strip()] < 2 and not _PAGE_NUMBER.match(l.strip())]
 757
 758    out: list[str] = []
 759    for line in kept:
 760        is_table_row = len(_COLUMNS.split(line)) > 1
 761        if out and not is_table_row and not out[-1].endswith("\n"):
 762            prev = out.pop()
 763            if prev.endswith("-") and line[:1].islower():
 764                out.append(prev[:-1] + line)  # hyphenated word: join with no space
 765            else:
 766                out.append(prev + " " + line)
 767        else:
 768            out.append(line + ("\n" if is_table_row else ""))
 769    return "\n".join(s.rstrip("\n") for s in out)
 770
 771
 772def extract_table(raw: str, header: list[str]) -> list[dict[str, str]]:
 773    """Find the table whose header row matches `header` and return its rows as dicts."""
 774    lines = [l.strip() for l in raw.split("\n")]
 775    for i, line in enumerate(lines):
 776        if _COLUMNS.split(line) == header:
 777            rows = []
 778            for row in lines[i + 1 :]:
 779                cells = _COLUMNS.split(row)
 780                if len(cells) != len(header):
 781                    break
 782                rows.append(dict(zip(header, cells)))
 783            return rows
 784    return []
 785
 786
 787def table_rows_as_sentences(rows: list[dict[str, str]]) -> list[str]:
 788    """One self-contained sentence per row, so each row can be retrieved on its own."""
 789    out = []
 790    for row in rows:
 791        key, *rest = list(row)
 792        out.append(f"{key} {row[key]}: " + ", ".join(f"{h.lower()} {row[h]}" for h in rest) + ".")
 793    return out
 794
 795
 796@dataclass(frozen=True)
 797class Passage:
 798    """A retrievable piece of a document, with the metadata filters need."""
 799
 800    id: str  # "<doc id>#<position>"
 801    doc_id: str
 802    title: str
 803    text: str
 804    department: str
 805    updated: str
 806    acl: frozenset[str]
 807    index_text: str = ""  # what BM25 and the embedder actually see
 808
 809
 810def _sentences(text: str) -> list[str]:
 811    return [s for s in re.split(r"(?<=[.!?])\s+", text.strip()) if s]
 812
 813
 814def chunk_document(doc: Doc, max_words: int = 15, contextual: bool = False) -> list[Passage]:
 815    """Split a document on sentence boundaries into passages of at most ~max_words words.
 816
 817    With `contextual=True`, each passage is indexed with a short line saying
 818    where it comes from (title and department), so a sentence like "Update
 819    AnyConnect and reboot" still says which error it fixes.
 820    """
 821    groups: list[list[str]] = []
 822    for s in _sentences(doc.text):
 823        if groups and len(" ".join(groups[-1] + [s]).split()) <= max_words:
 824            groups[-1].append(s)
 825        else:
 826            groups.append([s])
 827    out = []
 828    for i, g in enumerate(groups):
 829        text = " ".join(g)
 830        prefix = f"{doc.title} ({doc.department}): " if contextual else ""
 831        out.append(Passage(f"{doc.id}#{i}", doc.id, doc.title, text, doc.department, doc.updated, doc.acl, prefix + text))
 832    return out
 833
 834
 835ALL_GROUPS = frozenset(g for d in DOCS for g in d.acl)
 836
 837
 838class RAGIndex:
 839    """Passages from every document, indexed for keyword (BM25) and dense search."""
 840
 841    def __init__(self, docs: list[Doc] = DOCS, contextual: bool = False, embedder: ConceptEmbedder | None = None, max_words: int = 15):
 842        self.passages = [p for d in docs for p in chunk_document(d, max_words, contextual)]
 843        self.by_id = {p.id: p for p in self.passages}
 844        self.embedder = embedder or ConceptEmbedder()
 845        self.bm25 = BM25([tokenize(p.index_text) for p in self.passages])
 846        self.vectors = self.embedder.encode([p.index_text for p in self.passages])  # (n_passages, dim), unit rows
 847        self.reranker = CrossEncoder(self.embedder)
 848
 849    def retrieve(self, query: str, user_groups: set[str], k: int = 5, mode: str = "hybrid",
 850                 updated_after: str | None = None, depth: int = 10) -> list[Passage]:
 851        """Top-k passages the user may see, ranked by BM25, dense, or both fused.
 852
 853        The permission and date filters run FIRST, on the candidate set,
 854        before anything is ranked, so restricted text can't leak into the
 855        results, the context, or the answer.
 856        """
 857        allowed = [i for i, p in enumerate(self.passages)
 858                   if p.acl & set(user_groups) and (updated_after is None or p.updated >= updated_after)]
 859        rankings = []
 860        if mode in ("bm25", "hybrid"):
 861            s = self.bm25.scores(tokenize(query))
 862            rankings.append([self.passages[i].id for i in sorted(allowed, key=lambda i: -s[i]) if s[i] > 0][:depth])
 863        if mode in ("dense", "hybrid"):
 864            sims = self.vectors @ self.embedder.encode(query)
 865            rankings.append([self.passages[i].id for i in sorted(allowed, key=lambda i: -sims[i])][:depth])
 866        fused = rankings[0] if len(rankings) == 1 else [pid for pid, _ in reciprocal_rank_fusion(rankings)]
 867        return [self.by_id[pid] for pid in fused[:k]]
 868
 869    def rerank(self, query: str, passages: list[Passage], top_n: int = 3) -> list[Passage]:
 870        """Second stage: score each (question, passage) pair together; keep the best few."""
 871        as_docs = {p.id: Doc(p.id, p.title, p.text, p.department, p.updated, p.acl) for p in passages}
 872        return sorted(passages, key=lambda p: -self.reranker.score(query, as_docs[p.id]))[:top_n]
 873
 874    def retrieve_parents(self, query: str, user_groups: set[str], k: int = 3) -> list[Doc]:
 875        """Search small passages (precise matches), return their whole documents (full context)."""
 876        seen: list[str] = []
 877        for p in self.retrieve(query, user_groups, k=k * 4):
 878            if p.doc_id not in seen:
 879                seen.append(p.doc_id)
 880        return [DOCS_BY_ID[d] for d in seen[:k]]
 881
 882
 883# ===========================================================================
 884# 2. QUERY TIME: assemble context, generate, verify
 885# ===========================================================================
 886
 887SYSTEM = ("Answer only from the sources. Cite the id of each source you use in square brackets. "
 888          "If the sources don't contain the answer, say you don't know.")
 889DONT_KNOW = "I don't know based on the provided sources."
 890
 891
 892def assemble_context(question: str, passages: list[Passage]) -> str:
 893    """Sources in escaped, id-tagged blocks; the question last (see primer.agents.context)."""
 894    sources = "\n".join(xml_wrap("source", p.text, id=p.id, title=p.title, updated=p.updated) for p in passages)
 895    return f"<sources>\n{sources}\n</sources>\n<question>{question}</question>"
 896
 897
 898def _overlap(question: str, sentence: str) -> int:
 899    """Shared ideas between question and sentence; identifiers (with digits) count double."""
 900    shared = set(ideas(question)) & set(ideas(sentence))
 901    return sum(2 if any(c.isdigit() for c in i) else 1 for i in shared)
 902
 903
 904def grounded_policy(system: str, messages: list[dict], tools: list[dict] | None) -> str:  # noqa: ARG001
 905    """Offline stand-in for a model that answers only from tagged sources, with citations.
 906
 907    It picks the source sentence sharing the most ideas with the question
 908    and cites it; if nothing shares at least two, it declines. A real model
 909    writes more fluent answers, but the contract (answer from sources, cite,
 910    or decline) is the same one you'd put in its system prompt.
 911    """
 912    text = last_user_text(messages)
 913    question = re.search(r"<question>(.*?)</question>", text, re.S).group(1)
 914    best, best_score = None, 0
 915    for sid, body in re.findall(r'<source id="([^"]+)"[^>]*>(.*?)</source>', text, re.S):
 916        for sentence in _sentences(body):
 917            score = _overlap(question, sentence)
 918            if score > best_score:
 919                best, best_score = (sentence, sid), score
 920    if best is None or best_score < 2:
 921        return DONT_KNOW
 922    return f"{best[0]} [{best[1]}]"
 923
 924
 925def generate(question: str, passages: list[Passage], llm: Any = None) -> str:
 926    llm = llm or ScriptedLLM(grounded_policy)
 927    reply = llm.complete(system=SYSTEM, messages=[{"role": "user", "content": assemble_context(question, passages)}])
 928    return reply.text
 929
 930
 931_CITATION = re.compile(r"\[([\w-]+#\d+)\]")
 932_CLAIM_THEN_CITATION = re.compile(r"([^\[\]]+?)\s*\[([\w-]+#\d+)\]")
 933
 934
 935def verify_citations(answer_text: str, passages: list[Passage]) -> dict[str, list[str]]:
 936    """Check that every cited id was actually in the context and supports its sentence."""
 937    by_id = {p.id: p for p in passages}
 938    cited = _CITATION.findall(answer_text)
 939    unknown = [c for c in cited if c not in by_id]
 940    unsupported = []
 941    # Each citation vouches for the text between the previous citation and itself.
 942    for claim, c in _CLAIM_THEN_CITATION.findall(answer_text):
 943        claim = claim.strip(" .")
 944        if c in by_id and claim_support(claim, [by_id[c].text]) < 0.6:
 945            unsupported.append(claim)
 946    return {"cited": cited, "unknown_ids": unknown, "unsupported": unsupported}
 947
 948
 949@dataclass
 950class RAGResult:
 951    answer: str
 952    passages: list[Passage]
 953    context: str
 954    citations: dict[str, list[str]] = field(default_factory=dict)
 955
 956
 957def answer(index: RAGIndex, question: str, user_groups: set[str], k: int = 8, top_n: int = 3, llm: Any = None) -> RAGResult:
 958    """The whole query-time pipeline: retrieve (filtered) -> rerank -> assemble -> generate -> verify."""
 959    candidates = index.retrieve(question, user_groups, k=k)
 960    passages = index.rerank(question, candidates, top_n) if candidates else []
 961    text = generate(question, passages, llm)
 962    return RAGResult(text, passages, assemble_context(question, passages), verify_citations(text, passages))
 963
 964
 965def post_generation_filter_answer(index: RAGIndex, question: str, user_groups: set[str]) -> str:
 966    """The WRONG way: retrieve everything, generate, then strip restricted citations.
 967
 968    The citation markers go; the restricted facts the model already wrote stay.
 969    """
 970    result = answer(index, question, set(ALL_GROUPS))
 971    allowed = {p.id for p in index.passages if p.acl & set(user_groups)}
 972    return _CITATION.sub(lambda m: m.group(0) if m.group(1) in allowed else "", result.answer).strip()
 973
 974
 975# ===========================================================================
 976# 3. RETRIEVAL UPGRADES
 977# ===========================================================================
 978
 979_REFERRING = re.compile(r"\b(it|that|this|they|those|them|what about|and for)\b", re.I)
 980
 981
 982def rewrite_query(question: str, history: list[str], llm: Any = None) -> str:
 983    """Turn a follow-up into a standalone search query.
 984
 985    Deterministic stand-in: if the follow-up leans on earlier context ("it",
 986    "that", "what about"), append the content words of the previous user
 987    turn. In production a small model does this with an instruction like
 988    "rewrite the last question so it can be understood without the chat".
 989    """
 990    if llm is not None:
 991        prompt = "Rewrite the last question as a standalone search query.\n" + "\n".join(history + [question])
 992        return llm.complete(system="", messages=[{"role": "user", "content": prompt}]).text
 993    if not history or not _REFERRING.search(question):
 994        return question
 995    extra = [t for t in tokenize(history[-1]) if t not in tokenize(question)]
 996    return f"{question} {' '.join(extra)}"
 997
 998
 999# What a model might write as a hypothetical answer. HyDE only uses the draft
1000# to *search*; its details may be wrong, and it's never shown to the user.
1001HYDE_DRAFTS = {
1002    "I got a weird message asking for my bank details":
1003        "This sounds like a phishing email. Don't click links; report the suspicious email with the Report phishing button in Outlook.",
1004    "why can't my computer reach the office network from home?":
1005        "Remote access to internal systems needs the VPN. Install AnyConnect, sign in, and connect from home.",
1006}
1007
1008
1009def _draft_policy(system, messages, tools):  # noqa: ARG001
1010    q = last_user_text(messages)
1011    return HYDE_DRAFTS.get(q, q)
1012
1013
1014def hyde_search(index: RAGIndex, question: str, user_groups: set[str], k: int = 5, llm: Any = None) -> list[Passage]:
1015    """Hypothetical Document Embeddings: search with a drafted answer instead of the question.
1016
1017    Answers look like documents; questions don't. Embedding a plausible
1018    answer lands closer to the real passage than embedding the question.
1019    """
1020    llm = llm or ScriptedLLM(_draft_policy)
1021    draft = llm.complete(system="Write a short passage that answers the question.",
1022                         messages=[{"role": "user", "content": question}]).text
1023    return index.retrieve(draft, user_groups, k=k, mode="dense")
1024
1025
1026# Alternative phrasings a model might write for multi-query retrieval.
1027MULTI_QUERY_DRAFTS = {
1028    "two step login setup": ["two-factor authentication setup", "MFA enrollment authenticator app QR code"],
1029}
1030
1031
1032def _phrasings_policy(system, messages, tools):  # noqa: ARG001
1033    q = last_user_text(messages)
1034    return "\n".join(MULTI_QUERY_DRAFTS.get(q, []))
1035
1036
1037def multi_query_search(index: RAGIndex, question: str, user_groups: set[str], k: int = 5, llm: Any = None) -> list[Passage]:
1038    """Search several phrasings of the question and fuse the rankings with RRF."""
1039    llm = llm or ScriptedLLM(_phrasings_policy)
1040    reply = llm.complete(system="Write two alternative search queries, one per line.", messages=[{"role": "user", "content": question}])
1041    phrasings = [question] + [l for l in reply.text.split("\n") if l.strip()]
1042    rankings = [[p.id for p in index.retrieve(q, user_groups, k=10)] for q in phrasings]
1043    return [index.by_id[pid] for pid, _ in reciprocal_rank_fusion(rankings)[:k]]
1044
1045
1046# ===========================================================================
1047# 4. AGENTIC RAG: the model decides whether and what to search
1048# ===========================================================================
1049
1050SEARCH_TOOL = {"name": "search_kb", "description": "Search the company knowledge base. Returns passages with ids.",
1051               "input_schema": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]}}
1052_SMALL_TALK = re.compile(r"^\s*(thanks|thank you|hi|hello|ok|great|bye)\b", re.I)
1053
1054
1055def _agentic_policy(system, messages, tools):  # noqa: ARG001
1056    """Split multi-part questions, search once per part, then answer every part with citations."""
1057    question = last_user_text(messages)
1058    if _SMALL_TALK.match(question):
1059        return "You're welcome!"
1060    parts = [p.strip(" ?") for p in re.split(r"\band\b", question) if len(tokenize(p)) >= 2]
1061    done = tool_results(messages)
1062    if len(done) < len(parts):
1063        return ToolCall("", "search_kb", {"query": parts[len(done)]})
1064    answers = []
1065    for part, result in zip(parts, done):
1066        sub = grounded_policy("", [{"role": "user", "content": f"{result['content']}\n<question>{part}</question>"}], None)
1067        answers.append(sub)
1068    return " ".join(answers)
1069
1070
1071def agentic_answer(index: RAGIndex, question: str, user_groups: set[str], max_steps: int = 6, llm: Any = None) -> dict[str, Any]:
1072    llm = llm or ScriptedLLM(_agentic_policy)
1073    messages: list[dict] = [{"role": "user", "content": question}]
1074    searches: list[str] = []
1075    reply = None
1076    for _ in range(max_steps):
1077        reply = llm.complete(system=SYSTEM, messages=messages, tools=[SEARCH_TOOL])
1078        messages.append({"role": "assistant", "content": reply.assistant_content})
1079        if reply.stop_reason != "tool_use":
1080            break
1081        results = []
1082        for call in reply.tool_calls:
1083            q = call.input["query"]
1084            searches.append(q)
1085            # The tool applies the freshness filter, so superseded policies never come back.
1086            hits = index.rerank(q, index.retrieve(q, user_groups, k=8, updated_after="2024-01-01"), top_n=3)
1087            body = "<sources>" + "".join(xml_wrap("source", p.text, id=p.id) for p in hits) + "</sources>"
1088            results.append(tool_result_block(call.id, body))
1089        messages.append({"role": "user", "content": results})
1090    return {"answer": reply.text if reply else "", "searches": searches}
1091
1092
1093# ===========================================================================
1094# 5. GRAPHRAG, in miniature
1095# ===========================================================================
1096
1097ENTITIES: dict[str, str] = {
1098    "two-factor": r"two-factor|\bmfa\b",
1099    "VPN": r"\bvpn\b",
1100    "AnyConnect": r"anyconnect",
1101    "software center": r"software center",
1102    "authenticator app": r"authenticator app",
1103    "email": r"\bemails?\b",
1104    "Outlook": r"outlook",
1105    "travel portal": r"travel portal",
1106    "HR portal": r"hr portal",
1107    "IT portal": r"it portal",
1108    "expense tool": r"expense tool",
1109    "self-service portal": r"self-service portal",
1110    "payroll": r"payroll",
1111    "laptop": r"\blaptops?\b",
1112}
1113
1114
1115def build_graph(docs: list[Doc] = DOCS) -> dict[str, dict[str, set[str]]]:
1116    """Entity graph: two entities are linked when a sentence mentions both.
1117
1118    Each link remembers which documents stated it, so answers drawn from the
1119    graph can still cite sources. Real GraphRAG uses a model to extract
1120    entities and typed relations ("requires", "replaces") and then
1121    summarizes communities of related entities.
1122    """
1123    graph: dict[str, dict[str, set[str]]] = {}
1124    for d in docs:
1125        for sentence in _sentences(f"{d.title}. {d.text}"):
1126            found = sorted(name for name, pat in ENTITIES.items() if re.search(pat, sentence, re.I))
1127            for a in found:
1128                for b in found:
1129                    if a != b:
1130                        graph.setdefault(a, {}).setdefault(b, set()).add(d.id)
1131    return graph
1132
1133
1134def communities(graph: dict[str, dict[str, set[str]]]) -> list[set[str]]:
1135    """Connected groups of entities: the "themes" GraphRAG summarizes."""
1136    seen: set[str] = set()
1137    groups = []
1138    for start in graph:
1139        if start in seen:
1140            continue
1141        stack, comp = [start], set()
1142        while stack:
1143            n = stack.pop()
1144            if n not in comp:
1145                comp.add(n)
1146                stack.extend(graph.get(n, {}))
1147        seen |= comp
1148        groups.append(comp)
1149    return sorted(groups, key=len, reverse=True)
1150
1151
1152# ===========================================================================
1153# 6. DEBUGGING a confident wrong answer
1154# ===========================================================================
1155
1156
1157def debug_wrong_answers(index: RAGIndex, mode: str = "hybrid", k: int = 3, queries: list[tuple[str, set[str]]] = LABELED_QUERIES) -> dict[str, Any]:
1158    """Separate retrieval failures from generation failures on a labelled set.
1159
1160    For each question: if no relevant document is in the top k, it's a
1161    retrieval miss (no prompt change can fix it). If one is there but the
1162    answer doesn't cite it, it's a generation miss.
1163    """
1164    by_query: dict[str, str] = {}
1165    for q, relevant in queries:
1166        passages = index.retrieve(q, set(ALL_GROUPS), k=k, mode=mode)
1167        if not relevant & {p.doc_id for p in passages}:
1168            by_query[q] = "retrieval miss"
1169            continue
1170        text = generate(q, passages)
1171        cited_docs = {c.split("#")[0] for c in _CITATION.findall(text)}
1172        by_query[q] = "ok" if cited_docs & relevant else "generation miss"
1173    misses = sum(v == "retrieval miss" for v in by_query.values())
1174    return {"by_query": by_query, "recall_at_k": 1 - misses / len(queries)}
1175
1176
1177def recall_curve(index: RAGIndex, mode: str, ks: list[int], rerank: bool = False,
1178                 queries: list[tuple[str, set[str]]] = LABELED_QUERIES) -> list[float]:
1179    """recall@k for each k: share of questions with a relevant document in the top k."""
1180    out = []
1181    for k in ks:
1182        hits = 0
1183        for q, relevant in queries:
1184            if rerank:
1185                passages = index.rerank(q, index.retrieve(q, set(ALL_GROUPS), k=10, mode=mode), top_n=k)
1186            else:
1187                passages = index.retrieve(q, set(ALL_GROUPS), k=k, mode=mode)
1188            hits += bool(relevant & {p.doc_id for p in passages})
1189        out.append(hits / len(queries))
1190    return out
1191
1192
1193# ===========================================================================
1194# Figures and demo
1195# ===========================================================================
1196
1197
1198def figures() -> dict[str, Any]:
1199    import matplotlib
1200
1201    matplotlib.use("Agg")
1202    import matplotlib.pyplot as plt
1203
1204    figs: dict[str, Any] = {}
1205    index = RAGIndex()
1206    ks = [1, 2, 3, 4, 5]
1207    fig, ax = plt.subplots(figsize=(6.5, 3.6))
1208    for label, mode, rr, color in [("keyword (BM25)", "bm25", False, "#dd8452"), ("dense", "dense", False, "#8172b2"),
1209                                   ("hybrid (RRF)", "hybrid", False, "#4c72b0"), ("hybrid + rerank", "hybrid", True, "#55a868")]:
1210        ax.plot(ks, recall_curve(index, mode, ks, rr), "o-", label=label, color=color)
1211    ax.set_xticks(ks)
1212    ax.set_xlabel("k (passages kept)")
1213    ax.set_ylabel("recall@k on the labelled questions")
1214    ax.set_ylim(0, 1.05)
1215    ax.set_title("Which retrieval finds the answer?")
1216    ax.legend(fontsize=8, loc="lower right")
1217    fig.tight_layout()
1218    figs["recall"] = fig
1219
1220    fig, ax = plt.subplots(figsize=(6.5, 3.2))
1221    kinds = ["ok", "generation miss", "retrieval miss"]
1222    colors = {"ok": "#55a868", "generation miss": "#dd8452", "retrieval miss": "#c44e52"}
1223    modes = ["bm25", "dense", "hybrid"]
1224    left = np.zeros(len(modes))
1225    for kind in kinds:
1226        counts = np.array([list(debug_wrong_answers(index, m)["by_query"].values()).count(kind) for m in modes])
1227        ax.barh(modes, counts, left=left, color=colors[kind], label=kind)
1228        left += counts
1229    ax.set_xlabel("labelled questions (top 3 passages)")
1230    ax.set_title("Where do wrong answers come from?")
1231    ax.legend(fontsize=8, loc="lower right")
1232    fig.tight_layout()
1233    figs["diagnosis"] = fig
1234
1235    plain, ctx = RAGIndex(), RAGIndex(contextual=True)
1236    fix = "Update AnyConnect to version 5.1 or later and reboot."
1237    qs = ["how do I fix ERR-4012", "ERR-4012 what should I do", "resolve ERR-4012 on my laptop"]
1238
1239    def rank_of(idx, q):
1240        ids = [p.text for p in idx.retrieve(q, {"everyone"}, k=20)]
1241        return ids.index(fix) + 1 if fix in ids else 21
1242
1243    fig, ax = plt.subplots(figsize=(6.5, 3))
1244    y = np.arange(len(qs))
1245    ax.barh(y - 0.2, [rank_of(plain, q) for q in qs], 0.4, color="#c44e52", label="plain chunks")
1246    ax.barh(y + 0.2, [rank_of(ctx, q) for q in qs], 0.4, color="#4c72b0", label="chunks with a context line")
1247    ax.set_yticks(y, qs)
1248    ax.invert_yaxis()
1249    ax.set_xlabel("rank of the passage containing the fix (lower is better; 21 = not in top 20)")
1250    ax.set_title("Contextual retrieval")
1251    ax.legend(fontsize=8)
1252    fig.tight_layout()
1253    figs["contextual"] = fig
1254    return figs
1255
1256
1257def demo() -> None:
1258    index = RAGIndex()
1259    banner("1. Parsing: repair the extraction before anything else")
1260    print(clean_extracted_text(MESSY_PDF))
1261    print()
1262    for s in table_rows_as_sentences(extract_table(MESSY_PDF, ["Grade", "Hotel cap", "Meal cap"])):
1263        print("row ->", s)
1264    print()
1265
1266    banner("2. The pipeline, end to end")
1267    r = answer(index, "What does ERR-4012 mean?", {"everyone"})
1268    print(r.context)
1269    print()
1270    print("answer:", r.answer)
1271    print("citations:", r.citations)
1272    print()
1273
1274    banner("3. Permissions are filtered before retrieval, not after generation")
1275    print("finance user :", answer(index, "What is the Q3 revenue forecast?", {"everyone", "finance"}).answer)
1276    print("other user   :", answer(index, "What is the Q3 revenue forecast?", {"everyone"}).answer)
1277    print("post-filtered:", post_generation_filter_answer(index, "What is the Q3 revenue forecast?", {"everyone"}))
1278    print()
1279    say("Post-filtering removed the citation but the number had already been written. Filter first.")
1280
1281    banner("4. Retrieval upgrades")
1282    ctx = RAGIndex(contextual=True)
1283    rows = [
1284        ("hybrid vs dense on 'what does ERR-4012 mean'", [p.doc_id for p in index.retrieve("what does ERR-4012 mean", {"everyone"}, 3, "dense")], [p.doc_id for p in index.retrieve("what does ERR-4012 mean", {"everyone"}, 3)]),
1285        ("date filter on the travel policy", [p.doc_id for p in index.retrieve("business class flights over four hours", {"everyone"}, 3)], [p.doc_id for p in index.retrieve("business class flights over four hours", {"everyone"}, 3, updated_after="2024-01-01")]),
1286        ("contextual chunks for 'how do I fix ERR-4012'", [p.id for p in index.retrieve("how do I fix ERR-4012", {"everyone"}, 3)], [p.id for p in ctx.retrieve("how do I fix ERR-4012", {"everyone"}, 3)]),
1287        ("rewrite 'what if it fails?' after a VPN question", [p.doc_id for p in index.retrieve("what if it fails?", {"everyone"}, 3)], [p.doc_id for p in index.retrieve(rewrite_query("what if it fails?", ["How do I set up the VPN?"]), {"everyone"}, 3)]),
1288        ("HyDE for 'weird message asking for my bank details'", [p.doc_id for p in index.retrieve("I got a weird message asking for my bank details", {"everyone"}, 3, "dense")], [p.doc_id for p in hyde_search(index, "I got a weird message asking for my bank details", {"everyone"}, 3)]),
1289        ("multi-query for 'two step login setup'", [p.doc_id for p in index.retrieve("two step login setup", {"everyone"}, 3)], [p.doc_id for p in multi_query_search(index, "two step login setup", {"everyone"}, 3)]),
1290    ]
1291    table(["upgrade", "before", "after"], rows)
1292
1293    banner("5. Agentic RAG")
1294    print(agentic_answer(index, "What is the per-diem for meals and how do I upload expense receipts?", {"everyone"}))
1295    print(agentic_answer(index, "thanks, that's all!", {"everyone"}))
1296    print()
1297
1298    banner("6. GraphRAG in miniature")
1299    g = build_graph()
1300    for n, docs in sorted(g["two-factor"].items()):
1301        print(f"two-factor -- {n}  (stated in {sorted(docs)})")
1302    print("communities:", communities(g))
1303    print()
1304
1305    banner("7. Debugging confident wrong answers: retrieval or generation?")
1306    for mode in ("dense", "hybrid"):
1307        rep = debug_wrong_answers(index, mode)
1308        misses = {q: v for q, v in rep["by_query"].items() if v != "ok"}
1309        print(f"{mode:6s} recall@3 {rep['recall_at_k']:.2f}  problems: {misses}")
1310    print()
1311    takeaway("Measure retrieval first. If the right passage isn't in the top k, no prompt can fix the answer.")
1312
1313
1314if __name__ == "__main__":
1315    demo()
Level 3: the code, function by function.
MESSY_PDF = 'Northwind Ltd Travel Policy 2026 CONFIDENTIAL\nHotel and meal limits apply to all em-\nployees travelling on company busi-\nness. Book through the travel portal.\nGrade Hotel cap Meal cap\nL1-L3 150 60\nL4-L6 200 75\nPage 1 of 2\n\x0cNorthwind Ltd Travel Policy 2026 CONFIDENTIAL\nEconomy class is required for flights un-\nder six hours.\nPage 2 of 2\n'
def clean_extracted_text(raw: str) -> str: on GitHub
743def clean_extracted_text(raw: str) -> str:
744    """Repair the usual damage of PDF text extraction.
745
746    1. Drop lines repeated on two or more pages (running headers/footers).
747    2. Drop "Page n of m" footers.
748    3. Rejoin words hyphenated across line breaks ("em-" + "ployees").
749    4. Unwrap prose lines into paragraphs; keep table rows on their own lines.
750    """
751    pages = [p.strip("\n").split("\n") for p in raw.split("\f")]
752    seen_on: dict[str, int] = {}
753    for lines in pages:
754        for line in set(l.strip() for l in lines):
755            seen_on[line] = seen_on.get(line, 0) + 1
756    kept = [l.strip() for lines in pages for l in lines
757            if l.strip() and seen_on[l.strip()] < 2 and not _PAGE_NUMBER.match(l.strip())]
758
759    out: list[str] = []
760    for line in kept:
761        is_table_row = len(_COLUMNS.split(line)) > 1
762        if out and not is_table_row and not out[-1].endswith("\n"):
763            prev = out.pop()
764            if prev.endswith("-") and line[:1].islower():
765                out.append(prev[:-1] + line)  # hyphenated word: join with no space
766            else:
767                out.append(prev + " " + line)
768        else:
769            out.append(line + ("\n" if is_table_row else ""))
770    return "\n".join(s.rstrip("\n") for s in out)

Repair the usual damage of PDF text extraction.

  1. Drop lines repeated on two or more pages (running headers/footers).
  2. Drop "Page n of m" footers.
  3. Rejoin words hyphenated across line breaks ("em-" + "ployees").
  4. Unwrap prose lines into paragraphs; keep table rows on their own lines.
def extract_table(raw: str, header: list[str]) -> list[dict[str, str]]: on GitHub
773def extract_table(raw: str, header: list[str]) -> list[dict[str, str]]:
774    """Find the table whose header row matches `header` and return its rows as dicts."""
775    lines = [l.strip() for l in raw.split("\n")]
776    for i, line in enumerate(lines):
777        if _COLUMNS.split(line) == header:
778            rows = []
779            for row in lines[i + 1 :]:
780                cells = _COLUMNS.split(row)
781                if len(cells) != len(header):
782                    break
783                rows.append(dict(zip(header, cells)))
784            return rows
785    return []

Find the table whose header row matches header and return its rows as dicts.

def table_rows_as_sentences(rows: list[dict[str, str]]) -> list[str]: on GitHub
788def table_rows_as_sentences(rows: list[dict[str, str]]) -> list[str]:
789    """One self-contained sentence per row, so each row can be retrieved on its own."""
790    out = []
791    for row in rows:
792        key, *rest = list(row)
793        out.append(f"{key} {row[key]}: " + ", ".join(f"{h.lower()} {row[h]}" for h in rest) + ".")
794    return out

One self-contained sentence per row, so each row can be retrieved on its own.

@dataclass(frozen=True)
class Passage: on GitHub
797@dataclass(frozen=True)
798class Passage:
799    """A retrievable piece of a document, with the metadata filters need."""
800
801    id: str  # "<doc id>#<position>"
802    doc_id: str
803    title: str
804    text: str
805    department: str
806    updated: str
807    acl: frozenset[str]
808    index_text: str = ""  # what BM25 and the embedder actually see

A retrievable piece of a document, with the metadata filters need.

Passage( id: str, doc_id: str, title: str, text: str, department: str, updated: str, acl: frozenset[str], index_text: str = '')
id: str
doc_id: str
title: str
text: str
department: str
updated: str
acl: frozenset[str]
index_text: str = ''
def chunk_document( doc: primer.common.corpus.Doc, max_words: int = 15, contextual: bool = False) -> list[Passage]: on GitHub
815def chunk_document(doc: Doc, max_words: int = 15, contextual: bool = False) -> list[Passage]:
816    """Split a document on sentence boundaries into passages of at most ~max_words words.
817
818    With `contextual=True`, each passage is indexed with a short line saying
819    where it comes from (title and department), so a sentence like "Update
820    AnyConnect and reboot" still says which error it fixes.
821    """
822    groups: list[list[str]] = []
823    for s in _sentences(doc.text):
824        if groups and len(" ".join(groups[-1] + [s]).split()) <= max_words:
825            groups[-1].append(s)
826        else:
827            groups.append([s])
828    out = []
829    for i, g in enumerate(groups):
830        text = " ".join(g)
831        prefix = f"{doc.title} ({doc.department}): " if contextual else ""
832        out.append(Passage(f"{doc.id}#{i}", doc.id, doc.title, text, doc.department, doc.updated, doc.acl, prefix + text))
833    return out

Split a document on sentence boundaries into passages of at most ~max_words words.

With contextual=True, each passage is indexed with a short line saying where it comes from (title and department), so a sentence like "Update AnyConnect and reboot" still says which error it fixes.

ALL_GROUPS = frozenset({'exec', 'everyone', 'finance', 'hr'})
class RAGIndex: on GitHub
839class RAGIndex:
840    """Passages from every document, indexed for keyword (BM25) and dense search."""
841
842    def __init__(self, docs: list[Doc] = DOCS, contextual: bool = False, embedder: ConceptEmbedder | None = None, max_words: int = 15):
843        self.passages = [p for d in docs for p in chunk_document(d, max_words, contextual)]
844        self.by_id = {p.id: p for p in self.passages}
845        self.embedder = embedder or ConceptEmbedder()
846        self.bm25 = BM25([tokenize(p.index_text) for p in self.passages])
847        self.vectors = self.embedder.encode([p.index_text for p in self.passages])  # (n_passages, dim), unit rows
848        self.reranker = CrossEncoder(self.embedder)
849
850    def retrieve(self, query: str, user_groups: set[str], k: int = 5, mode: str = "hybrid",
851                 updated_after: str | None = None, depth: int = 10) -> list[Passage]:
852        """Top-k passages the user may see, ranked by BM25, dense, or both fused.
853
854        The permission and date filters run FIRST, on the candidate set,
855        before anything is ranked, so restricted text can't leak into the
856        results, the context, or the answer.
857        """
858        allowed = [i for i, p in enumerate(self.passages)
859                   if p.acl & set(user_groups) and (updated_after is None or p.updated >= updated_after)]
860        rankings = []
861        if mode in ("bm25", "hybrid"):
862            s = self.bm25.scores(tokenize(query))
863            rankings.append([self.passages[i].id for i in sorted(allowed, key=lambda i: -s[i]) if s[i] > 0][:depth])
864        if mode in ("dense", "hybrid"):
865            sims = self.vectors @ self.embedder.encode(query)
866            rankings.append([self.passages[i].id for i in sorted(allowed, key=lambda i: -sims[i])][:depth])
867        fused = rankings[0] if len(rankings) == 1 else [pid for pid, _ in reciprocal_rank_fusion(rankings)]
868        return [self.by_id[pid] for pid in fused[:k]]
869
870    def rerank(self, query: str, passages: list[Passage], top_n: int = 3) -> list[Passage]:
871        """Second stage: score each (question, passage) pair together; keep the best few."""
872        as_docs = {p.id: Doc(p.id, p.title, p.text, p.department, p.updated, p.acl) for p in passages}
873        return sorted(passages, key=lambda p: -self.reranker.score(query, as_docs[p.id]))[:top_n]
874
875    def retrieve_parents(self, query: str, user_groups: set[str], k: int = 3) -> list[Doc]:
876        """Search small passages (precise matches), return their whole documents (full context)."""
877        seen: list[str] = []
878        for p in self.retrieve(query, user_groups, k=k * 4):
879            if p.doc_id not in seen:
880                seen.append(p.doc_id)
881        return [DOCS_BY_ID[d] for d in seen[:k]]

Passages from every document, indexed for keyword (BM25) and dense search.

RAGIndex( docs: list[primer.common.corpus.Doc] = [Doc(id='it-001', title='How to reset your password', text="If you forgot your password or your account is locked, go to the self-service portal, choose 'Forgot password', verify with your authenticator app, and set a new password. The reset link expires after 15 minutes.", department='IT', updated='2026-03-02', acl=frozenset({'everyone'})), Doc(id='it-002', title='Password policy', text='Passwords must be at least 14 characters and include a number and a symbol. Passwords expire every 180 days and the last 10 passwords cannot be reused. This policy applies to all employees and contractors.', department='IT', updated='2025-11-20', acl=frozenset({'everyone'})), Doc(id='it-003', title='Setting up the VPN', text='Install the AnyConnect client from the software center, sign in with your company credentials, and approve the two-factor prompt. Use the VPN for all remote access to internal systems.', department='IT', updated='2026-01-15', acl=frozenset({'everyone'})), Doc(id='it-004', title='Error ERR-4012: VPN tunnel failed', text='ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. Update AnyConnect to version 5.1 or later and reboot.', department='IT', updated='2026-04-10', acl=frozenset({'everyone'})), Doc(id='it-005', title='Error ERR-4013: certificate expired', text="ERR-4013 means your device certificate expired. Open the software center and run 'Renew device certificate', then reconnect.", department='IT', updated='2026-04-10', acl=frozenset({'everyone'})), Doc(id='it-006', title='Setting up two-factor authentication', text='Download the authenticator app, scan the QR code on the security page, and enter the six digit code to finish enrolling in MFA. Two-factor is required for email and VPN.', department='IT', updated='2025-09-01', acl=frozenset({'everyone'})), Doc(id='it-007', title='Reporting phishing emails', text="If an email looks suspicious, do not click links. Use the 'Report phishing' button in Outlook. Security reviews every report within one business day.", department='IT', updated='2026-02-11', acl=frozenset({'everyone'})), Doc(id='it-008', title='Requesting a new laptop', text="Laptops are refreshed every three years. Submit a hardware request in the IT portal with your manager's approval. Standard devices are MacBook Pro or ThinkPad X1.", department='IT', updated='2025-12-05', acl=frozenset({'everyone'})), Doc(id='it-009', title='Printer troubleshooting', text='If printing fails, check the toner and paper tray, then remove and re-add the printer from settings. Floor printers are named by building and floor.', department='IT', updated='2024-06-30', acl=frozenset({'everyone'})), Doc(id='fin-001', title='Car mileage expense', text='Employees who use a personal vehicle for business driving are reimbursed at the standard mileage rate. Log each trip with date, distance and purpose, and submit the expense within 30 days.', department='Finance', updated='2026-01-05', acl=frozenset({'everyone'})), Doc(id='fin-002', title='Travel policy (2026)', text='Book flights and hotels through the travel portal. Economy class is required for flights under six hours. Meals are covered by a daily per-diem of 75 dollars.', department='Finance', updated='2026-01-01', acl=frozenset({'everyone'})), Doc(id='fin-003', title='Travel policy (2023, superseded)', text='Book flights through the travel agency by phone. Business class is allowed for flights over four hours. Meals are covered by a daily per-diem of 60 dollars.', department='Finance', updated='2023-01-01', acl=frozenset({'everyone'})), Doc(id='fin-004', title='Submitting expense receipts', text='Upload receipts to the expense tool within 30 days. Receipts are required for any expense over 25 dollars. Reimbursements are paid with the next payroll run.', department='Finance', updated='2025-10-12', acl=frozenset({'everyone'})), Doc(id='fin-005', title='Vendor invoice reconciliation', text='At quarter end, match each vendor invoice to its payment record. List every mismatch with the invoice number, amount and vendor, and send the reconciliation to the controller.', department='Finance', updated='2026-03-31', acl=frozenset({'finance'})), Doc(id='fin-006', title='Q3 revenue forecast', text='Q3 revenue is forecast at 41 million dollars, up 12 percent, driven by enterprise renewals.', department='Finance', updated='2026-07-01', acl=frozenset({'exec', 'finance'})), Doc(id='hr-001', title='Paid time off', text='Full-time employees accrue 20 days of PTO per year. Request vacation in the HR portal at least two weeks ahead. Unused PTO up to 5 days rolls over.', department='HR', updated='2026-01-01', acl=frozenset({'everyone'})), Doc(id='hr-002', title='Sick leave', text='Employees receive 10 paid sick days per year. No manager approval is needed for sick leave, but notify your team as early as possible.', department='HR', updated='2025-08-15', acl=frozenset({'everyone'})), Doc(id='hr-003', title='New hire onboarding', text='On your first day, collect your laptop from IT, complete security training, and set up two-factor authentication. Your manager will schedule orientation sessions for week one.', department='HR', updated='2026-02-01', acl=frozenset({'everyone'})), Doc(id='hr-004', title='Salary bands and bonus', text='Salary bands are reviewed every April. The annual bonus target is 10 percent of base pay, paid in March based on company and individual performance.', department='HR', updated='2026-04-01', acl=frozenset({'exec', 'hr'})), Doc(id='hr-005', title='Parental leave', text='Primary caregivers receive 16 weeks of paid parental leave; secondary caregivers receive 6 weeks. Leave can start up to two weeks before the expected birth or adoption date.', department='HR', updated='2025-05-20', acl=frozenset({'everyone'}))], contextual: bool = False, embedder: primer.common.embedder.ConceptEmbedder | None = None, max_words: int = 15) on GitHub
842    def __init__(self, docs: list[Doc] = DOCS, contextual: bool = False, embedder: ConceptEmbedder | None = None, max_words: int = 15):
843        self.passages = [p for d in docs for p in chunk_document(d, max_words, contextual)]
844        self.by_id = {p.id: p for p in self.passages}
845        self.embedder = embedder or ConceptEmbedder()
846        self.bm25 = BM25([tokenize(p.index_text) for p in self.passages])
847        self.vectors = self.embedder.encode([p.index_text for p in self.passages])  # (n_passages, dim), unit rows
848        self.reranker = CrossEncoder(self.embedder)
passages
embedder
bm25
vectors
reranker
def retrieve( self, query: str, user_groups: set[str], k: int = 5, mode: str = 'hybrid', updated_after: str | None = None, depth: int = 10) -> list[Passage]: on GitHub
850    def retrieve(self, query: str, user_groups: set[str], k: int = 5, mode: str = "hybrid",
851                 updated_after: str | None = None, depth: int = 10) -> list[Passage]:
852        """Top-k passages the user may see, ranked by BM25, dense, or both fused.
853
854        The permission and date filters run FIRST, on the candidate set,
855        before anything is ranked, so restricted text can't leak into the
856        results, the context, or the answer.
857        """
858        allowed = [i for i, p in enumerate(self.passages)
859                   if p.acl & set(user_groups) and (updated_after is None or p.updated >= updated_after)]
860        rankings = []
861        if mode in ("bm25", "hybrid"):
862            s = self.bm25.scores(tokenize(query))
863            rankings.append([self.passages[i].id for i in sorted(allowed, key=lambda i: -s[i]) if s[i] > 0][:depth])
864        if mode in ("dense", "hybrid"):
865            sims = self.vectors @ self.embedder.encode(query)
866            rankings.append([self.passages[i].id for i in sorted(allowed, key=lambda i: -sims[i])][:depth])
867        fused = rankings[0] if len(rankings) == 1 else [pid for pid, _ in reciprocal_rank_fusion(rankings)]
868        return [self.by_id[pid] for pid in fused[:k]]

Top-k passages the user may see, ranked by BM25, dense, or both fused.

The permission and date filters run FIRST, on the candidate set, before anything is ranked, so restricted text can't leak into the results, the context, or the answer.

def rerank( self, query: str, passages: list[Passage], top_n: int = 3) -> list[Passage]: on GitHub
870    def rerank(self, query: str, passages: list[Passage], top_n: int = 3) -> list[Passage]:
871        """Second stage: score each (question, passage) pair together; keep the best few."""
872        as_docs = {p.id: Doc(p.id, p.title, p.text, p.department, p.updated, p.acl) for p in passages}
873        return sorted(passages, key=lambda p: -self.reranker.score(query, as_docs[p.id]))[:top_n]

Second stage: score each (question, passage) pair together; keep the best few.

def retrieve_parents( self, query: str, user_groups: set[str], k: int = 3) -> list[primer.common.corpus.Doc]: on GitHub
875    def retrieve_parents(self, query: str, user_groups: set[str], k: int = 3) -> list[Doc]:
876        """Search small passages (precise matches), return their whole documents (full context)."""
877        seen: list[str] = []
878        for p in self.retrieve(query, user_groups, k=k * 4):
879            if p.doc_id not in seen:
880                seen.append(p.doc_id)
881        return [DOCS_BY_ID[d] for d in seen[:k]]

Search small passages (precise matches), return their whole documents (full context).

SYSTEM = "Answer only from the sources. Cite the id of each source you use in square brackets. If the sources don't contain the answer, say you don't know."
DONT_KNOW = "I don't know based on the provided sources."
def assemble_context(question: str, passages: list[Passage]) -> str: on GitHub
893def assemble_context(question: str, passages: list[Passage]) -> str:
894    """Sources in escaped, id-tagged blocks; the question last (see primer.agents.context)."""
895    sources = "\n".join(xml_wrap("source", p.text, id=p.id, title=p.title, updated=p.updated) for p in passages)
896    return f"<sources>\n{sources}\n</sources>\n<question>{question}</question>"

Sources in escaped, id-tagged blocks; the question last (see primer.agents.context).

def grounded_policy(system: str, messages: list[dict], tools: list[dict] | None) -> str: on GitHub
905def grounded_policy(system: str, messages: list[dict], tools: list[dict] | None) -> str:  # noqa: ARG001
906    """Offline stand-in for a model that answers only from tagged sources, with citations.
907
908    It picks the source sentence sharing the most ideas with the question
909    and cites it; if nothing shares at least two, it declines. A real model
910    writes more fluent answers, but the contract (answer from sources, cite,
911    or decline) is the same one you'd put in its system prompt.
912    """
913    text = last_user_text(messages)
914    question = re.search(r"<question>(.*?)</question>", text, re.S).group(1)
915    best, best_score = None, 0
916    for sid, body in re.findall(r'<source id="([^"]+)"[^>]*>(.*?)</source>', text, re.S):
917        for sentence in _sentences(body):
918            score = _overlap(question, sentence)
919            if score > best_score:
920                best, best_score = (sentence, sid), score
921    if best is None or best_score < 2:
922        return DONT_KNOW
923    return f"{best[0]} [{best[1]}]"

Offline stand-in for a model that answers only from tagged sources, with citations.

It picks the source sentence sharing the most ideas with the question and cites it; if nothing shares at least two, it declines. A real model writes more fluent answers, but the contract (answer from sources, cite, or decline) is the same one you'd put in its system prompt.

def generate( question: str, passages: list[Passage], llm: Any = None) -> str: on GitHub
926def generate(question: str, passages: list[Passage], llm: Any = None) -> str:
927    llm = llm or ScriptedLLM(grounded_policy)
928    reply = llm.complete(system=SYSTEM, messages=[{"role": "user", "content": assemble_context(question, passages)}])
929    return reply.text
def verify_citations( answer_text: str, passages: list[Passage]) -> dict[str, list[str]]: on GitHub
936def verify_citations(answer_text: str, passages: list[Passage]) -> dict[str, list[str]]:
937    """Check that every cited id was actually in the context and supports its sentence."""
938    by_id = {p.id: p for p in passages}
939    cited = _CITATION.findall(answer_text)
940    unknown = [c for c in cited if c not in by_id]
941    unsupported = []
942    # Each citation vouches for the text between the previous citation and itself.
943    for claim, c in _CLAIM_THEN_CITATION.findall(answer_text):
944        claim = claim.strip(" .")
945        if c in by_id and claim_support(claim, [by_id[c].text]) < 0.6:
946            unsupported.append(claim)
947    return {"cited": cited, "unknown_ids": unknown, "unsupported": unsupported}

Check that every cited id was actually in the context and supports its sentence.

@dataclass
class RAGResult: on GitHub
950@dataclass
951class RAGResult:
952    answer: str
953    passages: list[Passage]
954    context: str
955    citations: dict[str, list[str]] = field(default_factory=dict)
RAGResult( answer: str, passages: list[Passage], context: str, citations: dict[str, list[str]] = <factory>)
answer: str
passages: list[Passage]
context: str
citations: dict[str, list[str]]
def answer( index: RAGIndex, question: str, user_groups: set[str], k: int = 8, top_n: int = 3, llm: Any = None) -> RAGResult: on GitHub
958def answer(index: RAGIndex, question: str, user_groups: set[str], k: int = 8, top_n: int = 3, llm: Any = None) -> RAGResult:
959    """The whole query-time pipeline: retrieve (filtered) -> rerank -> assemble -> generate -> verify."""
960    candidates = index.retrieve(question, user_groups, k=k)
961    passages = index.rerank(question, candidates, top_n) if candidates else []
962    text = generate(question, passages, llm)
963    return RAGResult(text, passages, assemble_context(question, passages), verify_citations(text, passages))

The whole query-time pipeline: retrieve (filtered) -> rerank -> assemble -> generate -> verify.

def post_generation_filter_answer( index: RAGIndex, question: str, user_groups: set[str]) -> str: on GitHub
966def post_generation_filter_answer(index: RAGIndex, question: str, user_groups: set[str]) -> str:
967    """The WRONG way: retrieve everything, generate, then strip restricted citations.
968
969    The citation markers go; the restricted facts the model already wrote stay.
970    """
971    result = answer(index, question, set(ALL_GROUPS))
972    allowed = {p.id for p in index.passages if p.acl & set(user_groups)}
973    return _CITATION.sub(lambda m: m.group(0) if m.group(1) in allowed else "", result.answer).strip()

The WRONG way: retrieve everything, generate, then strip restricted citations.

The citation markers go; the restricted facts the model already wrote stay.

def rewrite_query(question: str, history: list[str], llm: Any = None) -> str: on GitHub
983def rewrite_query(question: str, history: list[str], llm: Any = None) -> str:
984    """Turn a follow-up into a standalone search query.
985
986    Deterministic stand-in: if the follow-up leans on earlier context ("it",
987    "that", "what about"), append the content words of the previous user
988    turn. In production a small model does this with an instruction like
989    "rewrite the last question so it can be understood without the chat".
990    """
991    if llm is not None:
992        prompt = "Rewrite the last question as a standalone search query.\n" + "\n".join(history + [question])
993        return llm.complete(system="", messages=[{"role": "user", "content": prompt}]).text
994    if not history or not _REFERRING.search(question):
995        return question
996    extra = [t for t in tokenize(history[-1]) if t not in tokenize(question)]
997    return f"{question} {' '.join(extra)}"

Turn a follow-up into a standalone search query.

Deterministic stand-in: if the follow-up leans on earlier context ("it", "that", "what about"), append the content words of the previous user turn. In production a small model does this with an instruction like "rewrite the last question so it can be understood without the chat".

HYDE_DRAFTS = {'I got a weird message asking for my bank details': "This sounds like a phishing email. Don't click links; report the suspicious email with the Report phishing button in Outlook.", "why can't my computer reach the office network from home?": 'Remote access to internal systems needs the VPN. Install AnyConnect, sign in, and connect from home.'}
MULTI_QUERY_DRAFTS = {'two step login setup': ['two-factor authentication setup', 'MFA enrollment authenticator app QR code']}
SEARCH_TOOL = {'name': 'search_kb', 'description': 'Search the company knowledge base. Returns passages with ids.', 'input_schema': {'type': 'object', 'properties': {'query': {'type': 'string'}}, 'required': ['query']}}
def agentic_answer( index: RAGIndex, question: str, user_groups: set[str], max_steps: int = 6, llm: Any = None) -> dict[str, typing.Any]: on GitHub
1072def agentic_answer(index: RAGIndex, question: str, user_groups: set[str], max_steps: int = 6, llm: Any = None) -> dict[str, Any]:
1073    llm = llm or ScriptedLLM(_agentic_policy)
1074    messages: list[dict] = [{"role": "user", "content": question}]
1075    searches: list[str] = []
1076    reply = None
1077    for _ in range(max_steps):
1078        reply = llm.complete(system=SYSTEM, messages=messages, tools=[SEARCH_TOOL])
1079        messages.append({"role": "assistant", "content": reply.assistant_content})
1080        if reply.stop_reason != "tool_use":
1081            break
1082        results = []
1083        for call in reply.tool_calls:
1084            q = call.input["query"]
1085            searches.append(q)
1086            # The tool applies the freshness filter, so superseded policies never come back.
1087            hits = index.rerank(q, index.retrieve(q, user_groups, k=8, updated_after="2024-01-01"), top_n=3)
1088            body = "<sources>" + "".join(xml_wrap("source", p.text, id=p.id) for p in hits) + "</sources>"
1089            results.append(tool_result_block(call.id, body))
1090        messages.append({"role": "user", "content": results})
1091    return {"answer": reply.text if reply else "", "searches": searches}
ENTITIES: dict[str, str] = {'two-factor': 'two-factor|\\bmfa\\b', 'VPN': '\\bvpn\\b', 'AnyConnect': 'anyconnect', 'software center': 'software center', 'authenticator app': 'authenticator app', 'email': '\\bemails?\\b', 'Outlook': 'outlook', 'travel portal': 'travel portal', 'HR portal': 'hr portal', 'IT portal': 'it portal', 'expense tool': 'expense tool', 'self-service portal': 'self-service portal', 'payroll': 'payroll', 'laptop': '\\blaptops?\\b'}
def build_graph( docs: list[primer.common.corpus.Doc] = [Doc(id='it-001', title='How to reset your password', text="If you forgot your password or your account is locked, go to the self-service portal, choose 'Forgot password', verify with your authenticator app, and set a new password. The reset link expires after 15 minutes.", department='IT', updated='2026-03-02', acl=frozenset({'everyone'})), Doc(id='it-002', title='Password policy', text='Passwords must be at least 14 characters and include a number and a symbol. Passwords expire every 180 days and the last 10 passwords cannot be reused. This policy applies to all employees and contractors.', department='IT', updated='2025-11-20', acl=frozenset({'everyone'})), Doc(id='it-003', title='Setting up the VPN', text='Install the AnyConnect client from the software center, sign in with your company credentials, and approve the two-factor prompt. Use the VPN for all remote access to internal systems.', department='IT', updated='2026-01-15', acl=frozenset({'everyone'})), Doc(id='it-004', title='Error ERR-4012: VPN tunnel failed', text='ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. Update AnyConnect to version 5.1 or later and reboot.', department='IT', updated='2026-04-10', acl=frozenset({'everyone'})), Doc(id='it-005', title='Error ERR-4013: certificate expired', text="ERR-4013 means your device certificate expired. Open the software center and run 'Renew device certificate', then reconnect.", department='IT', updated='2026-04-10', acl=frozenset({'everyone'})), Doc(id='it-006', title='Setting up two-factor authentication', text='Download the authenticator app, scan the QR code on the security page, and enter the six digit code to finish enrolling in MFA. Two-factor is required for email and VPN.', department='IT', updated='2025-09-01', acl=frozenset({'everyone'})), Doc(id='it-007', title='Reporting phishing emails', text="If an email looks suspicious, do not click links. Use the 'Report phishing' button in Outlook. Security reviews every report within one business day.", department='IT', updated='2026-02-11', acl=frozenset({'everyone'})), Doc(id='it-008', title='Requesting a new laptop', text="Laptops are refreshed every three years. Submit a hardware request in the IT portal with your manager's approval. Standard devices are MacBook Pro or ThinkPad X1.", department='IT', updated='2025-12-05', acl=frozenset({'everyone'})), Doc(id='it-009', title='Printer troubleshooting', text='If printing fails, check the toner and paper tray, then remove and re-add the printer from settings. Floor printers are named by building and floor.', department='IT', updated='2024-06-30', acl=frozenset({'everyone'})), Doc(id='fin-001', title='Car mileage expense', text='Employees who use a personal vehicle for business driving are reimbursed at the standard mileage rate. Log each trip with date, distance and purpose, and submit the expense within 30 days.', department='Finance', updated='2026-01-05', acl=frozenset({'everyone'})), Doc(id='fin-002', title='Travel policy (2026)', text='Book flights and hotels through the travel portal. Economy class is required for flights under six hours. Meals are covered by a daily per-diem of 75 dollars.', department='Finance', updated='2026-01-01', acl=frozenset({'everyone'})), Doc(id='fin-003', title='Travel policy (2023, superseded)', text='Book flights through the travel agency by phone. Business class is allowed for flights over four hours. Meals are covered by a daily per-diem of 60 dollars.', department='Finance', updated='2023-01-01', acl=frozenset({'everyone'})), Doc(id='fin-004', title='Submitting expense receipts', text='Upload receipts to the expense tool within 30 days. Receipts are required for any expense over 25 dollars. Reimbursements are paid with the next payroll run.', department='Finance', updated='2025-10-12', acl=frozenset({'everyone'})), Doc(id='fin-005', title='Vendor invoice reconciliation', text='At quarter end, match each vendor invoice to its payment record. List every mismatch with the invoice number, amount and vendor, and send the reconciliation to the controller.', department='Finance', updated='2026-03-31', acl=frozenset({'finance'})), Doc(id='fin-006', title='Q3 revenue forecast', text='Q3 revenue is forecast at 41 million dollars, up 12 percent, driven by enterprise renewals.', department='Finance', updated='2026-07-01', acl=frozenset({'exec', 'finance'})), Doc(id='hr-001', title='Paid time off', text='Full-time employees accrue 20 days of PTO per year. Request vacation in the HR portal at least two weeks ahead. Unused PTO up to 5 days rolls over.', department='HR', updated='2026-01-01', acl=frozenset({'everyone'})), Doc(id='hr-002', title='Sick leave', text='Employees receive 10 paid sick days per year. No manager approval is needed for sick leave, but notify your team as early as possible.', department='HR', updated='2025-08-15', acl=frozenset({'everyone'})), Doc(id='hr-003', title='New hire onboarding', text='On your first day, collect your laptop from IT, complete security training, and set up two-factor authentication. Your manager will schedule orientation sessions for week one.', department='HR', updated='2026-02-01', acl=frozenset({'everyone'})), Doc(id='hr-004', title='Salary bands and bonus', text='Salary bands are reviewed every April. The annual bonus target is 10 percent of base pay, paid in March based on company and individual performance.', department='HR', updated='2026-04-01', acl=frozenset({'exec', 'hr'})), Doc(id='hr-005', title='Parental leave', text='Primary caregivers receive 16 weeks of paid parental leave; secondary caregivers receive 6 weeks. Leave can start up to two weeks before the expected birth or adoption date.', department='HR', updated='2025-05-20', acl=frozenset({'everyone'}))]) -> dict[str, dict[str, set[str]]]: on GitHub
1116def build_graph(docs: list[Doc] = DOCS) -> dict[str, dict[str, set[str]]]:
1117    """Entity graph: two entities are linked when a sentence mentions both.
1118
1119    Each link remembers which documents stated it, so answers drawn from the
1120    graph can still cite sources. Real GraphRAG uses a model to extract
1121    entities and typed relations ("requires", "replaces") and then
1122    summarizes communities of related entities.
1123    """
1124    graph: dict[str, dict[str, set[str]]] = {}
1125    for d in docs:
1126        for sentence in _sentences(f"{d.title}. {d.text}"):
1127            found = sorted(name for name, pat in ENTITIES.items() if re.search(pat, sentence, re.I))
1128            for a in found:
1129                for b in found:
1130                    if a != b:
1131                        graph.setdefault(a, {}).setdefault(b, set()).add(d.id)
1132    return graph

Entity graph: two entities are linked when a sentence mentions both.

Each link remembers which documents stated it, so answers drawn from the graph can still cite sources. Real GraphRAG uses a model to extract entities and typed relations ("requires", "replaces") and then summarizes communities of related entities.

def communities(graph: dict[str, dict[str, set[str]]]) -> list[set[str]]: on GitHub
1135def communities(graph: dict[str, dict[str, set[str]]]) -> list[set[str]]:
1136    """Connected groups of entities: the "themes" GraphRAG summarizes."""
1137    seen: set[str] = set()
1138    groups = []
1139    for start in graph:
1140        if start in seen:
1141            continue
1142        stack, comp = [start], set()
1143        while stack:
1144            n = stack.pop()
1145            if n not in comp:
1146                comp.add(n)
1147                stack.extend(graph.get(n, {}))
1148        seen |= comp
1149        groups.append(comp)
1150    return sorted(groups, key=len, reverse=True)

Connected groups of entities: the "themes" GraphRAG summarizes.

def debug_wrong_answers( index: RAGIndex, mode: str = 'hybrid', k: int = 3, queries: list[tuple[str, set[str]]] = [("I forgot my password and I'm locked out", {'it-001'}), ('how long must a password be', {'it-002'}), ('what does ERR-4012 mean', {'it-004'}), ('ERR-4013 on my laptop', {'it-005'}), ('automobile reimbursement for business driving', {'fin-001'}), ('how many vacation days do I get', {'hr-001'}), ('connect to internal systems from home', {'it-003'}), ('suspicious email with a link', {'it-007'}), ('per-diem for meals when traveling', {'fin-002'}), ('enroll in MFA', {'it-006'}), ('first day checklist for a new hire', {'hr-003'}), ('match vendor bills to payments', {'fin-005'})]) -> dict[str, typing.Any]: on GitHub
1158def debug_wrong_answers(index: RAGIndex, mode: str = "hybrid", k: int = 3, queries: list[tuple[str, set[str]]] = LABELED_QUERIES) -> dict[str, Any]:
1159    """Separate retrieval failures from generation failures on a labelled set.
1160
1161    For each question: if no relevant document is in the top k, it's a
1162    retrieval miss (no prompt change can fix it). If one is there but the
1163    answer doesn't cite it, it's a generation miss.
1164    """
1165    by_query: dict[str, str] = {}
1166    for q, relevant in queries:
1167        passages = index.retrieve(q, set(ALL_GROUPS), k=k, mode=mode)
1168        if not relevant & {p.doc_id for p in passages}:
1169            by_query[q] = "retrieval miss"
1170            continue
1171        text = generate(q, passages)
1172        cited_docs = {c.split("#")[0] for c in _CITATION.findall(text)}
1173        by_query[q] = "ok" if cited_docs & relevant else "generation miss"
1174    misses = sum(v == "retrieval miss" for v in by_query.values())
1175    return {"by_query": by_query, "recall_at_k": 1 - misses / len(queries)}

Separate retrieval failures from generation failures on a labelled set.

For each question: if no relevant document is in the top k, it's a retrieval miss (no prompt change can fix it). If one is there but the answer doesn't cite it, it's a generation miss.

def recall_curve( index: RAGIndex, mode: str, ks: list[int], rerank: bool = False, queries: list[tuple[str, set[str]]] = [("I forgot my password and I'm locked out", {'it-001'}), ('how long must a password be', {'it-002'}), ('what does ERR-4012 mean', {'it-004'}), ('ERR-4013 on my laptop', {'it-005'}), ('automobile reimbursement for business driving', {'fin-001'}), ('how many vacation days do I get', {'hr-001'}), ('connect to internal systems from home', {'it-003'}), ('suspicious email with a link', {'it-007'}), ('per-diem for meals when traveling', {'fin-002'}), ('enroll in MFA', {'it-006'}), ('first day checklist for a new hire', {'hr-003'}), ('match vendor bills to payments', {'fin-005'})]) -> list[float]: on GitHub
1178def recall_curve(index: RAGIndex, mode: str, ks: list[int], rerank: bool = False,
1179                 queries: list[tuple[str, set[str]]] = LABELED_QUERIES) -> list[float]:
1180    """recall@k for each k: share of questions with a relevant document in the top k."""
1181    out = []
1182    for k in ks:
1183        hits = 0
1184        for q, relevant in queries:
1185            if rerank:
1186                passages = index.rerank(q, index.retrieve(q, set(ALL_GROUPS), k=10, mode=mode), top_n=k)
1187            else:
1188                passages = index.retrieve(q, set(ALL_GROUPS), k=k, mode=mode)
1189            hits += bool(relevant & {p.doc_id for p in passages})
1190        out.append(hits / len(queries))
1191    return out

recall@k for each k: share of questions with a relevant document in the top k.

def figures() -> dict[str, typing.Any]: on GitHub
1199def figures() -> dict[str, Any]:
1200    import matplotlib
1201
1202    matplotlib.use("Agg")
1203    import matplotlib.pyplot as plt
1204
1205    figs: dict[str, Any] = {}
1206    index = RAGIndex()
1207    ks = [1, 2, 3, 4, 5]
1208    fig, ax = plt.subplots(figsize=(6.5, 3.6))
1209    for label, mode, rr, color in [("keyword (BM25)", "bm25", False, "#dd8452"), ("dense", "dense", False, "#8172b2"),
1210                                   ("hybrid (RRF)", "hybrid", False, "#4c72b0"), ("hybrid + rerank", "hybrid", True, "#55a868")]:
1211        ax.plot(ks, recall_curve(index, mode, ks, rr), "o-", label=label, color=color)
1212    ax.set_xticks(ks)
1213    ax.set_xlabel("k (passages kept)")
1214    ax.set_ylabel("recall@k on the labelled questions")
1215    ax.set_ylim(0, 1.05)
1216    ax.set_title("Which retrieval finds the answer?")
1217    ax.legend(fontsize=8, loc="lower right")
1218    fig.tight_layout()
1219    figs["recall"] = fig
1220
1221    fig, ax = plt.subplots(figsize=(6.5, 3.2))
1222    kinds = ["ok", "generation miss", "retrieval miss"]
1223    colors = {"ok": "#55a868", "generation miss": "#dd8452", "retrieval miss": "#c44e52"}
1224    modes = ["bm25", "dense", "hybrid"]
1225    left = np.zeros(len(modes))
1226    for kind in kinds:
1227        counts = np.array([list(debug_wrong_answers(index, m)["by_query"].values()).count(kind) for m in modes])
1228        ax.barh(modes, counts, left=left, color=colors[kind], label=kind)
1229        left += counts
1230    ax.set_xlabel("labelled questions (top 3 passages)")
1231    ax.set_title("Where do wrong answers come from?")
1232    ax.legend(fontsize=8, loc="lower right")
1233    fig.tight_layout()
1234    figs["diagnosis"] = fig
1235
1236    plain, ctx = RAGIndex(), RAGIndex(contextual=True)
1237    fix = "Update AnyConnect to version 5.1 or later and reboot."
1238    qs = ["how do I fix ERR-4012", "ERR-4012 what should I do", "resolve ERR-4012 on my laptop"]
1239
1240    def rank_of(idx, q):
1241        ids = [p.text for p in idx.retrieve(q, {"everyone"}, k=20)]
1242        return ids.index(fix) + 1 if fix in ids else 21
1243
1244    fig, ax = plt.subplots(figsize=(6.5, 3))
1245    y = np.arange(len(qs))
1246    ax.barh(y - 0.2, [rank_of(plain, q) for q in qs], 0.4, color="#c44e52", label="plain chunks")
1247    ax.barh(y + 0.2, [rank_of(ctx, q) for q in qs], 0.4, color="#4c72b0", label="chunks with a context line")
1248    ax.set_yticks(y, qs)
1249    ax.invert_yaxis()
1250    ax.set_xlabel("rank of the passage containing the fix (lower is better; 21 = not in top 20)")
1251    ax.set_title("Contextual retrieval")
1252    ax.legend(fontsize=8)
1253    fig.tight_layout()
1254    figs["contextual"] = fig
1255    return figs
def demo() -> None: on GitHub
1258def demo() -> None:
1259    index = RAGIndex()
1260    banner("1. Parsing: repair the extraction before anything else")
1261    print(clean_extracted_text(MESSY_PDF))
1262    print()
1263    for s in table_rows_as_sentences(extract_table(MESSY_PDF, ["Grade", "Hotel cap", "Meal cap"])):
1264        print("row ->", s)
1265    print()
1266
1267    banner("2. The pipeline, end to end")
1268    r = answer(index, "What does ERR-4012 mean?", {"everyone"})
1269    print(r.context)
1270    print()
1271    print("answer:", r.answer)
1272    print("citations:", r.citations)
1273    print()
1274
1275    banner("3. Permissions are filtered before retrieval, not after generation")
1276    print("finance user :", answer(index, "What is the Q3 revenue forecast?", {"everyone", "finance"}).answer)
1277    print("other user   :", answer(index, "What is the Q3 revenue forecast?", {"everyone"}).answer)
1278    print("post-filtered:", post_generation_filter_answer(index, "What is the Q3 revenue forecast?", {"everyone"}))
1279    print()
1280    say("Post-filtering removed the citation but the number had already been written. Filter first.")
1281
1282    banner("4. Retrieval upgrades")
1283    ctx = RAGIndex(contextual=True)
1284    rows = [
1285        ("hybrid vs dense on 'what does ERR-4012 mean'", [p.doc_id for p in index.retrieve("what does ERR-4012 mean", {"everyone"}, 3, "dense")], [p.doc_id for p in index.retrieve("what does ERR-4012 mean", {"everyone"}, 3)]),
1286        ("date filter on the travel policy", [p.doc_id for p in index.retrieve("business class flights over four hours", {"everyone"}, 3)], [p.doc_id for p in index.retrieve("business class flights over four hours", {"everyone"}, 3, updated_after="2024-01-01")]),
1287        ("contextual chunks for 'how do I fix ERR-4012'", [p.id for p in index.retrieve("how do I fix ERR-4012", {"everyone"}, 3)], [p.id for p in ctx.retrieve("how do I fix ERR-4012", {"everyone"}, 3)]),
1288        ("rewrite 'what if it fails?' after a VPN question", [p.doc_id for p in index.retrieve("what if it fails?", {"everyone"}, 3)], [p.doc_id for p in index.retrieve(rewrite_query("what if it fails?", ["How do I set up the VPN?"]), {"everyone"}, 3)]),
1289        ("HyDE for 'weird message asking for my bank details'", [p.doc_id for p in index.retrieve("I got a weird message asking for my bank details", {"everyone"}, 3, "dense")], [p.doc_id for p in hyde_search(index, "I got a weird message asking for my bank details", {"everyone"}, 3)]),
1290        ("multi-query for 'two step login setup'", [p.doc_id for p in index.retrieve("two step login setup", {"everyone"}, 3)], [p.doc_id for p in multi_query_search(index, "two step login setup", {"everyone"}, 3)]),
1291    ]
1292    table(["upgrade", "before", "after"], rows)
1293
1294    banner("5. Agentic RAG")
1295    print(agentic_answer(index, "What is the per-diem for meals and how do I upload expense receipts?", {"everyone"}))
1296    print(agentic_answer(index, "thanks, that's all!", {"everyone"}))
1297    print()
1298
1299    banner("6. GraphRAG in miniature")
1300    g = build_graph()
1301    for n, docs in sorted(g["two-factor"].items()):
1302        print(f"two-factor -- {n}  (stated in {sorted(docs)})")
1303    print("communities:", communities(g))
1304    print()
1305
1306    banner("7. Debugging confident wrong answers: retrieval or generation?")
1307    for mode in ("dense", "hybrid"):
1308        rep = debug_wrong_answers(index, mode)
1309        misses = {q: v for q, v in rep["by_query"].items() if v != "ok"}
1310        print(f"{mode:6s} recall@3 {rep['recall_at_k']:.2f}  problems: {misses}")
1311    print()
1312    takeaway("Measure retrieval first. If the right passage isn't in the top k, no prompt can fix the answer.")