primer.ml.embeddings.operations
Operating embeddings: model changes, jargon, and finding what broke
Run: python -m primer.ml.embeddings.operations
New to vectors or Σ? primer.notation builds them from zero. This lesson
builds on primer.ml.embeddings.similarity and primer.ml.embeddings.contrastive.
Level 1: The practitioner's guide
In one sentence. Operating embeddings means keeping every query on the same map as the documents it searches, changing that map without a bad day, checking that the map speaks your domain's language, and knowing which half of a retrieval-augmented system failed when an answer is wrong.
When you need it. From the day an embedding index serves real traffic. Every embedding model is its own coordinate system: in this lesson, the word "vpn" embedded by two versions of the same toy model, both 64 dimensions, has a cosine similarity of about 0.05, as if they were unrelated words. Search documents embedded by one model with queries embedded by another and recall@3 on this lesson's golden set falls from 0.92 to 0.25, barely above the 0.15 that random ranking would score, and nothing crashes. You will change models: new releases, a fine-tune, a chunking change and a bug fix all require re-embedding everything. You will meet jargon the model doesn't know: this lesson's general model puts the right document first for one of four company-jargon questions. And you will get a confident wrong answer and need to know whether retrieval or generation caused it. The tell: "we upgraded the embedding model and search got weird", or weeks spent tuning prompts for answers whose right page was never retrieved.
Your options. Two decisions recur: how to change the model, and what to do when it doesn't know your words.
| Option | What it does | What it gives you | What it costs | Where it lives |
|---|---|---|---|---|
| Re-embed only new documents | Old documents keep their old vectors; new ones get the new model | Nothing: it is the classic incident. Queries and documents stop sharing a space and recall collapses silently | The cheapest to run and the most expensive to discover | A pipeline that forgot the backfill |
| Re-embed in place | Stops writes, re-embeds every document into the same index, resumes | A correct index at the end | Downtime or a window of mixed results, and no rollback except re-embedding again | Your ingestion job |
| Blue/green migration | Builds a second index alongside, dual-writes new documents to both, backfills, compares on a golden set, then repoints a live alias | No downtime, a gate that refuses a worse model, rollback by flipping one pointer | Double storage during the migration, a golden set, the backfill time | Your code plus index aliases (Elasticsearch, Qdrant) |
| Keep the general model, add keyword search | Fuses BM25 with dense search so exact jargon matches by spelling | Identifiers and acronyms found without training | Two searches per query | Any search engine (primer.ml.embeddings.retrieval) |
| Fine-tune on domain pairs | Trains the embedding model on your own (question, document) pairs | Four of four jargon questions right in this lesson, against one of four | Labeled pairs, a training run, then a full re-embed | primer.ml.embeddings.contrastive, followed by a migration |
| Measure and triage | Keeps a golden set of questions with known answer passages and checks it on every change | Silent failures made visible, each split into retrieval or generation | A few hundred labeled questions, taken from resolved logs | Your evaluation harness; Ragas |
How to choose.
- Changing the model on a live index: blue/green, every time. Version the index by model, dual-write from the first minute so nothing falls behind, backfill in restartable batches, and cut over only when the new index is at least as good on your golden set. Keep the old index for a soak period.
- Estimating the migration: multiply documents by tokens per document. Fifty million documents of 500 tokens is 25 billion tokens: about 7 hours at a million tokens per second across your workers, and \$500 at \$0.02 per million tokens (replace all three inputs with your own). The money is usually modest; the time, the rate limits and the double storage are what need planning.
- The model doesn't know your jargon: measure on your own questions first, because public leaderboards measure public data. Add keyword search for exact terms, then fine-tune on domain pairs if the gap remains.
- A wrong answer: ask one question first. Was a relevant document in the retrieved top k? If not, it is retrieval: chunking, hybrid search, reranking, the model. If yes and the answer ignored it, it is generation: the prompt, the context order, fewer distracting chunks.
- Whatever you do, keep the golden set and rerun it on every change. It is the only number that predicts your system's quality.
What it costs. A golden set: a few hundred questions with their answer documents, taken from resolved support or search sessions, which cost nothing to collect. Re-embedding: linear in the corpus, so ten times the documents means ten times the hours and the dollars, plus double storage while both indexes exist. Blue/green adds one extra write per new document during the migration and one shadow evaluation. Fine-tuning adds labeled pairs and a training run, and then a migration, because a fine-tuned model is a new map. Triage costs a person reading a handful of failures; only the retrieval half can be measured without a model, which is why recall@k on the golden set is the first number to track.
What breaks.
- Mixed spaces. Documents from one model, queries from another, the same number of dimensions: recall falls from 0.92 to 0.25 with no error. A model with a different number of dimensions at least crashes. Tie every index to a model version and refuse vectors from any other.
- A backfill that never finished. Only new documents got re-embedded. Dual-writing and backfilling are separate steps, and the migration is not done until every old document has been re-embedded.
- Cutting over on faith. The shadow comparison is the gate: this lesson's migration refuses a candidate that scores 0.58 against the live index's 0.92.
- No way back. Retire the old index only after a soak period; until then, rollback is repointing the alias.
- Trusting a leaderboard. BEIR showed that retrieval models trained on one domain often lose badly on others, and MTEB found no single model wins everywhere. Your questions are the benchmark.
- Tuning the prompt for a retrieval failure. The error-code question in this lesson stays wrong at every k because its page ranks fourth; no prompt fixes that. Retrieving more turns some retrieval failures into generation failures, so raising k is not a fix either.
In the wild. Index aliases are the switch: Elasticsearch's aliases API swaps an alias from one index to another in a single atomic operation with no downtime, and Qdrant's collection aliases let you build a second collection in the background and switch atomically with no concurrent request affected. Qdrant also fixes the vector size per collection, so a model with a new number of dimensions is a new collection by construction. Martin Fowler's description of blue/green deployment is where the pattern's name comes from. The MTEB leaderboard on Hugging Face ranks embedding models across tasks, and BEIR is the zero-shot retrieval benchmark that showed how much domain matters. Ragas separates retrieval metrics (context precision, context recall) from generation metrics (faithfulness, response relevancy), the same split as this lesson's triage. The papers behind this lesson are listed at the end.
Go deeper. Level 2 measures the mixed-space failure on the golden set, walks a blue/green migration state by state with the code that refuses a worse model, works the re-embedding arithmetic, shows the jargon gap and a tuned model closing it, and runs every golden question through a triage that names the failing half. If you only needed the runbook, you are done.
Level 2: How it works, from scratch
Level 2 makes each of those truths concrete, starting with two mapmakers.
The everyday picture. Two mapmakers each draw a map of the same city with their own grid. On one map, square (3, 7) is the train station; on the other, (3, 7) is a park. Both maps are fine, but you can't read a position off one map and look it up on the other.
Every embedding model is its own mapmaker. Upgrade the model and every document has to be re-plotted on the new map, and until that's done, the old map and the new one must never be mixed. That's the first of three operational truths this lesson makes concrete:
- Changing models means re-embedding everything, and doing it without downtime or a silent quality drop takes a plan.
- General models don't speak your company's jargon, so measure on your own questions and adapt when needed.
- When a retrieval-augmented system answers wrongly, find out which half failed: retrieval (the right page was never found) or generation (it was found and misused).
Every model has its own space
Tiny worked example: the word "vpn" embedded by two versions of the toy
model (primer.common.embedder) with the same 64 dimensions: their cosine
is about 0.06, essentially unrelated, though it's the same word. Now take 12
real questions with known answers (the "golden set" in primer.common.corpus)
over 20 documents. Search v1 documents with v1 queries and the right answer
is in the top 3 for 11 of 12 questions (0.92). Search the same v1
documents with queries from the new v2 model and it drops to 0.25, about
what random ranking would give (3 of 20 documents shown, so 15%), and
nothing crashes.
Level 3: the formula and its symbols
$$ \text{hit@}k = \frac{1}{|Q|} \sum_{q \in Q} \big[\, \text{rel}(q) \cap \text{top}_k(q) \neq \varnothing \,\big] $$
Symbols
| Symbol | Meaning here | Example |
|---|---|---|
| Q | the golden set of questions | 12 questions |
| |Q| | how many questions | 12 |
| q ∈ Q | each question in turn | |
| rel(q) | the documents that truly answer q | {it-004} |
| top_k(q) | the k documents search returned | 3 documents |
| ∩, ≠ ∅ | "they share at least one document" | |
| [ … ] | 1 if true, 0 if false | |
| hit@k | the share of questions with a right answer in the top k | 0 to 1 |
In words: the share of golden questions for which at least one right document appears among the top k results. (With one relevant document per question, as here, this equals recall@k.)
On the example: 11 hits out of 12 = 0.92 with matching models; 3 out of 12 = 0.25 with mixed models.
Level 3: in Python
In Python:
def hit(rel_q, top_q):
# [rel(q) ∩ top_k(q) ≠ ∅]
return 1 if rel_q & top_q else 0
hit({"it-004"}, {"it-002", "it-004", "fin-001"}) # → 1
# matching models: 11 of the 12 questions
hits = [1] * 11 + [0]
# (1/|Q|) Σ over q in Q
round(sum(hits) / len(hits), 2) # → 0.92
# mixed models: 3 of the 12
hits = [1] * 3 + [0] * 9
round(sum(hits) / len(hits), 2) # → 0.25
Reading it: each bar is recall@3 on the golden set. The first two bars use one model for both documents and queries, v1 then v2: both work. The third bar searches v1 documents with v2 queries: recall falls to 0.25 (3 of the 12 questions), barely above the dashed line at 0.15, which is what random ranking would score (3 of 20 documents shown). A model with a different number of dimensions would at least crash (you can't dot a 128-number vector with a 64-number one); a same-size model fails silently, which is worse.
In code: golden_recall embeds the documents with one model and the
questions with another and returns hit@k. VectorIndex is a flat index tied
to one model version: VectorIndex.search embeds a text query, and
VectorIndex.search_vector takes a vector that is already made.
Why it matters: "we upgraded the embedding model and search got weird" is a classic incident. The cause is almost always documents and queries embedded by different models, typically because only new documents were re-embedded.
Migrating without a bad day: blue/green
Everyday picture: building a new bridge next to the old one. Traffic keeps using the old bridge while the new one is built and inspected. When the new bridge passes inspection, traffic is switched over, and the old bridge stays standing for a while in case something turns up.
stateDiagram-v2 [*] --> Live_v1 Live_v1 --> DualWrite: start v2 (empty index, new docs go to both) DualWrite --> Backfilled: backfill (re-embed old docs in batches) Backfilled --> Compared: shadow compare (recall on golden set) Compared --> Live_v2: cutover, only if v2 at least as good Compared --> Live_v1: refused, v2 is worse Live_v2 --> Live_v1: rollback (v1 index kept) Live_v2 --> [*]: retire v1 after a soak period
Reading it: follow the states from the top. Users are served by v1 the
whole time until cutover. Dual-writing from the very first step means
documents added mid-migration land in both indexes, so the new one never
falls behind. The shadow comparison is the gate: the switch only happens
if v2 is at least as good on your own golden set (BlueGreenMigration
refuses to cut over without it). The old index is kept after cutover, so
rollback is flipping one pointer, not a multi-hour rebuild.
The "live alias" is a pointer: the search service asks for "the live index", and cutover or rollback just repoints it. Many vector databases support aliases for exactly this.
In code: each arrow of the diagram is one method: BlueGreenMigration.start,
BlueGreenMigration.add (the dual write), BlueGreenMigration.backfill,
BlueGreenMigration.shadow_compare, BlueGreenMigration.cutover and
BlueGreenMigration.rollback.
Tiny worked example: what will re-embedding cost? 50 million documents of about 500 tokens each, an embedding throughput of 1 million tokens per second across your workers, and a price of \$0.02 per million tokens (all three are inputs you replace with your own numbers):
Level 3: the formula and its symbols
$$ \text{hours} = \frac{n \cdot t}{r \cdot 3600}, \qquad \text{dollars} = \frac{n \cdot t}{10^6} \cdot p $$
Symbols
| Symbol | Meaning here | Example |
|---|---|---|
| n | number of documents (or chunks) | 50,000,000 |
| t | average tokens per document | 500 |
| n · t | total tokens to embed | 25,000,000,000 |
| r | throughput, tokens per second | 1,000,000 |
| 3600 | seconds per hour | |
| p | price per million tokens | \$0.02 |
In words: total tokens divided by throughput gives the time; total tokens in millions times the price gives the cost.
On the example: 25 × 10⁹ / (10⁶ × 3600) = 6.94 hours; 25,000 × \$0.02 = \$500.
Level 3: in Python
In Python:
# documents, tokens per document
n, t = 50_000_000, 500
# tokens per second, dollars per million tokens
r, p = 1_000_000, 0.02
# total tokens
n * t # → 25000000000
# hours
round(n * t / (r * 3600), 2) # → 6.94
# dollars
round(n * t / 10**6 * p, 2) # → 500.0
Reading it: the horizontal axis is corpus size (log scale); the left panel shows hours at three throughputs, the right panel shows dollars at three prices. Both grow in straight lines on these log axes: ten times the documents, ten times the time and money. The money is usually modest; the time, the rate limits and the double storage during the migration are what need planning.
In code: reembed_estimate computes the tokens, hours and dollars for
your own n, t, r and p.
Why it matters: re-embedding is routine: new models, fine-tunes, chunking changes and bug fixes all require it. Versioned indexes, dual writes, a golden-set gate and a rollback path turn it from a risky event into a boring one.
Domain mismatch: fluent, but not in your jargon
Everyday picture: a new hire who speaks perfect English but doesn't yet know that "AP" means accounts payable or that "T&E" means travel and expenses. General embedding models are trained mostly on web text; your acronyms and product names are new words to them.
Tiny worked example: four real-sounding employee questions in company
jargon: "ap aging report", "hotspot from a client site", "sso lockout",
"t-and-e submission deadline". The general toy model puts the right
document first for 1 of 4. A version that has learned the four jargon
words (DomainTunedEmbedder, standing in for a model fine-tuned on
company pairs) gets 4 of 4.
Reading it: each pair of bars is one jargon question; the height is where the right document ranked (1 is best, shorter is better). The general model buries three of the four answers; the tuned model puts every one first. The one the general model gets right, "sso lockout", is saved by a word it does know ("lockout").
In code: rank_of_first_relevant finds where the right document lands
for one question under one model: the height of each bar.
flowchart LR L["Search & support logs"] --> F["Keep resolved sessions<br/>(the clicked doc answered it)"] F --> N["Normalize and merge<br/>repeated questions"] N --> G["Golden set: question → relevant docs<br/>(a few hundred is plenty)"] G --> T["Test several models on YOUR set"] T --> D{Good enough?} D -->|no| FT["Fine-tune on domain pairs,<br/>add hybrid search"] D -->|yes| S[Ship, and keep the set for regressions] FT --> T
Reading it: the evaluation set comes from real usage, not invention.
Resolved sessions give you (question, answer document) pairs for free;
build_eval_set keeps only resolved ones and merges repeats. Test
candidate models on that set, and when none is good enough, fine-tune on
domain pairs (primer.ml.embeddings.contrastive) and add keyword search,
which matches jargon exactly (primer.ml.embeddings.retrieval).
Why it matters: public leaderboards measure public data. The only number that predicts your system's quality is recall on your own questions.
Measure retrieval separately from generation
Everyday picture: an open-book exam. A wrong answer has one of two causes: the right page wasn't in the book you brought (retrieval failure), or it was, and you misread it (generation failure). Studying harder fixes the second, not the first.
Tiny worked example: three answered questions, checked by hand.
| Retrieved (top 3) | Truly relevant | Answer used | Verdict |
|---|---|---|---|
| it-002, fin-006, fin-001 | it-004 | it-002 | retrieval failure: it-004 was never found |
| hr-002, hr-001, hr-005 | hr-001 | hr-002 | generation failure: found but not used |
| it-001, it-002 | it-001 | it-001 | correct |
flowchart TD A[Wrong answer] --> R{Was a relevant document<br/>in the retrieved top k?} R -->|no| RF[Retrieval failure:<br/>fix chunking, hybrid search,<br/>reranking, the embedding model] R -->|yes| G{Did the answer use it?} G -->|no| GF[Generation failure:<br/>fix the prompt, context order,<br/>fewer distracting chunks] G -->|yes| OK[Check the grader:<br/>the answer may be fine]
Reading it: one question splits every failure in two. If the right document wasn't retrieved, no prompt change can help, so work on retrieval. If it was retrieved and ignored, the problem is downstream, in the prompt or the context. Only the retrieval half can be measured without a model, which is why recall@k on a golden set is the first number to track.
Reading it: each bar is the 12 golden questions, answered by a toy generator that always uses the top document, with k documents retrieved. Green is correct, red is a retrieval failure, orange is a generation failure. Retrieving more (larger k) turns some retrieval failures into generation failures: the right page is now in the book, but the reader still opened the wrong one. Watch the error-code question ("what does ERR-4012 mean"): its answer ranks only fourth, so it's a retrieval failure until k = 5, and even then the top document is the wrong one. Dense vectors blur exact identifiers; keyword search would put that page first.
In code: triage gives the verdict for one answered question, following
the flowchart above; triage_golden_set runs every golden question through
retrieval and a toy generator that answers from the top document.
Why it matters: teams burn weeks tuning prompts for failures that were retrieval all along. Triage first, then fix the half that's broken.
In 20 seconds
- Every embedding model has its own vector space; never mix documents and queries from different models. Same-size models fail silently.
- Migrate blue/green: new index, dual-write, backfill, compare on a golden set, cut over only if better, keep the old index for rollback.
- Re-embedding cost is n × tokens: estimate hours and dollars before you start.
- General models miss company jargon; build an eval set from resolved logs, and fine-tune or add keyword search when needed.
- Triage RAG failures into retrieval vs. generation before fixing anything.
Self-test questions
Q: You need to switch embedding models on a 50-million-document index. Plan the migration. Create a versioned index for the new model and dual-write new documents to both. Backfill by re-embedding existing documents in restartable batches (estimate: 50M × 500 tokens = 25B tokens; at 1M tokens/s that's about 7 hours). Compare recall on a golden set in shadow; cut over the live alias only if the new model is at least as good; keep the old index for fast rollback, then retire it.
Q: Why can't you compare vectors from two different embedding models? Each model defines its own coordinate system. Even with the same number of dimensions, the same text lands in unrelated places, so similarity between the two spaces is meaningless.
Q: A general embedding model performs poorly on a company's internal documents. What helps? Build a small eval set from real queries and the documents that resolved them, test several models on it, add hybrid (keyword + dense) search for exact terms, and fine-tune an embedding model on domain pairs with hard negatives if needed.
Q: Your RAG system gives confident wrong answers. What's your first diagnostic step? Split retrieval from generation: for failing questions, check whether a relevant document was in the retrieved top k. If not, it's retrieval (fix search); if yes, it's generation (fix prompt and context). Track recall@k on a labeled set continuously.
The papers behind this lesson
- Thakur et al., BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models (2021): https://arxiv.org/abs/2104.08663. Showed that retrieval models trained on one domain often lose badly on others, which is why you evaluate on your own data.
- Muennighoff et al., MTEB: Massive Text Embedding Benchmark (2022): https://arxiv.org/abs/2210.07316. Compared embedding models across many tasks and found no single model wins everywhere.
Further reading
- Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023): https://arxiv.org/abs/2309.15217
- Martin Fowler, BlueGreenDeployment: https://martinfowler.com/bliki/BlueGreenDeployment.html
- MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard
1r""" 2# Operating embeddings: model changes, jargon, and finding what broke 3 4Run: `python -m primer.ml.embeddings.operations` 5 6New to vectors or Σ? `primer.notation` builds them from zero. This lesson 7builds on `primer.ml.embeddings.similarity` and `primer.ml.embeddings.contrastive`. 8 9## Level 1: The practitioner's guide 10 11**In one sentence.** Operating embeddings means keeping every query on the 12same map as the documents it searches, changing that map without a bad day, 13checking that the map speaks your domain's language, and knowing which half 14of a retrieval-augmented system failed when an answer is wrong. 15 16**When you need it.** From the day an embedding index serves real traffic. 17Every embedding model is its own coordinate system: in this lesson, the word 18"vpn" embedded by two versions of the same toy model, both 64 dimensions, 19has a cosine similarity of about 0.05, as if they were unrelated words. 20Search documents embedded by one model with queries embedded by another and 21recall@3 on this lesson's golden set falls from 0.92 to 0.25, barely above 22the 0.15 that random ranking would score, and nothing crashes. You will 23change models: new releases, a fine-tune, a chunking change and a bug fix 24all require re-embedding everything. You will meet jargon the model doesn't 25know: this lesson's general model puts the right document first for one of 26four company-jargon questions. And you will get a confident wrong answer and 27need to know whether retrieval or generation caused it. The tell: "we 28upgraded the embedding model and search got weird", or weeks spent tuning 29prompts for answers whose right page was never retrieved. 30 31**Your options.** Two decisions recur: how to change the model, and what to 32do when it doesn't know your words. 33 34| Option | What it does | What it gives you | What it costs | Where it lives | 35|---|---|---|---|---| 36| Re-embed only new documents | Old documents keep their old vectors; new ones get the new model | Nothing: it is the classic incident. Queries and documents stop sharing a space and recall collapses silently | The cheapest to run and the most expensive to discover | A pipeline that forgot the backfill | 37| Re-embed in place | Stops writes, re-embeds every document into the same index, resumes | A correct index at the end | Downtime or a window of mixed results, and no rollback except re-embedding again | Your ingestion job | 38| Blue/green migration | Builds a second index alongside, dual-writes new documents to both, backfills, compares on a golden set, then repoints a live alias | No downtime, a gate that refuses a worse model, rollback by flipping one pointer | Double storage during the migration, a golden set, the backfill time | Your code plus index aliases (Elasticsearch, Qdrant) | 39| Keep the general model, add keyword search | Fuses BM25 with dense search so exact jargon matches by spelling | Identifiers and acronyms found without training | Two searches per query | Any search engine (`primer.ml.embeddings.retrieval`) | 40| Fine-tune on domain pairs | Trains the embedding model on your own (question, document) pairs | Four of four jargon questions right in this lesson, against one of four | Labeled pairs, a training run, then a full re-embed | `primer.ml.embeddings.contrastive`, followed by a migration | 41| Measure and triage | Keeps a golden set of questions with known answer passages and checks it on every change | Silent failures made visible, each split into retrieval or generation | A few hundred labeled questions, taken from resolved logs | Your evaluation harness; Ragas | 42 43**How to choose.** 44 45- Changing the model on a live index: blue/green, every time. Version the 46 index by model, dual-write from the first minute so nothing falls behind, 47 backfill in restartable batches, and cut over only when the new index is 48 at least as good on your golden set. Keep the old index for a soak period. 49- Estimating the migration: multiply documents by tokens per document. 50 Fifty million documents of 500 tokens is 25 billion tokens: about 7 hours 51 at a million tokens per second across your workers, and \$500 at \$0.02 52 per million tokens (replace all three inputs with your own). The money is 53 usually modest; the time, the rate limits and the double storage are what 54 need planning. 55- The model doesn't know your jargon: measure on your own questions first, 56 because public leaderboards measure public data. Add keyword search for 57 exact terms, then fine-tune on domain pairs if the gap remains. 58- A wrong answer: ask one question first. Was a relevant document in the 59 retrieved top k? If not, it is retrieval: chunking, hybrid search, 60 reranking, the model. If yes and the answer ignored it, it is generation: 61 the prompt, the context order, fewer distracting chunks. 62- Whatever you do, keep the golden set and rerun it on every change. It is 63 the only number that predicts your system's quality. 64 65**What it costs.** A golden set: a few hundred questions with their answer 66documents, taken from resolved support or search sessions, which cost 67nothing to collect. Re-embedding: linear in the corpus, so ten times the 68documents means ten times the hours and the dollars, plus double storage 69while both indexes exist. Blue/green adds one extra write per new document 70during the migration and one shadow evaluation. Fine-tuning adds labeled 71pairs and a training run, and then a migration, because a fine-tuned model 72is a new map. Triage costs a person reading a handful of failures; only the 73retrieval half can be measured without a model, which is why recall@k on 74the golden set is the first number to track. 75 76**What breaks.** 77 78- **Mixed spaces.** Documents from one model, queries from another, the same 79 number of dimensions: recall falls from 0.92 to 0.25 with no error. A 80 model with a different number of dimensions at least crashes. Tie every 81 index to a model version and refuse vectors from any other. 82- **A backfill that never finished.** Only new documents got re-embedded. 83 Dual-writing and backfilling are separate steps, and the migration is not 84 done until every old document has been re-embedded. 85- **Cutting over on faith.** The shadow comparison is the gate: this 86 lesson's migration refuses a candidate that scores 0.58 against the live 87 index's 0.92. 88- **No way back.** Retire the old index only after a soak period; until 89 then, rollback is repointing the alias. 90- **Trusting a leaderboard.** BEIR showed that retrieval models trained on 91 one domain often lose badly on others, and MTEB found no single model 92 wins everywhere. Your questions are the benchmark. 93- **Tuning the prompt for a retrieval failure.** The error-code question in 94 this lesson stays wrong at every k because its page ranks fourth; no 95 prompt fixes that. Retrieving more turns some retrieval failures into 96 generation failures, so raising k is not a fix either. 97 98**In the wild.** Index aliases are the switch: Elasticsearch's aliases API 99swaps an alias from one index to another in a single atomic operation with 100no downtime, and Qdrant's collection aliases let you build a second 101collection in the background and switch atomically with no concurrent 102request affected. Qdrant also fixes the vector size per collection, so a 103model with a new number of dimensions is a new collection by construction. 104Martin Fowler's description of blue/green deployment is where the pattern's 105name comes from. The MTEB leaderboard on Hugging Face ranks embedding models 106across tasks, and BEIR is the zero-shot retrieval benchmark that showed how 107much domain matters. Ragas separates retrieval metrics (context precision, 108context recall) from generation metrics (faithfulness, response relevancy), 109the same split as this lesson's triage. The papers behind this lesson are 110listed at the end. 111 112**Go deeper.** Level 2 measures the mixed-space failure on the golden set, 113walks a blue/green migration state by state with the code that refuses a 114worse model, works the re-embedding arithmetic, shows the jargon gap and a 115tuned model closing it, and runs every golden question through a triage 116that names the failing half. If you only needed the runbook, you are done. 117 118## Level 2: How it works, from scratch 119 120Level 2 makes each of those truths concrete, starting with two mapmakers. 121 122**The everyday picture.** Two mapmakers each draw a map of the same city with their own grid. On one 123map, square (3, 7) is the train station; on the other, (3, 7) is a park. Both 124maps are fine, but you can't read a position off one map and look it up on 125the other. 126 127Every embedding model is its own mapmaker. Upgrade the model and every 128document has to be re-plotted on the new map, and until that's done, the old 129map and the new one must never be mixed. That's the first of three 130operational truths this lesson makes concrete: 131 1321. **Changing models means re-embedding everything**, and doing it without 133 downtime or a silent quality drop takes a plan. 1342. **General models don't speak your company's jargon**, so measure on your 135 own questions and adapt when needed. 1363. **When a retrieval-augmented system answers wrongly, find out which half 137 failed**: retrieval (the right page was never found) or generation (it was 138 found and misused). 139 140## Every model has its own space 141 142**Tiny worked example:** the word "vpn" embedded by two versions of the toy 143model (`primer.common.embedder`) with the same 64 dimensions: their cosine 144is about 0.06, essentially unrelated, though it's the same word. Now take 12 145real questions with known answers (the "golden set" in `primer.common.corpus`) 146over 20 documents. Search v1 documents with v1 queries and the right answer 147is in the top 3 for 11 of 12 questions (**0.92**). Search the *same* v1 148documents with queries from the new v2 model and it drops to **0.25**, about 149what random ranking would give (3 of 20 documents shown, so 15%), and 150**nothing crashes**. 151 152$$ 153\text{hit@}k = \frac{1}{|Q|} \sum_{q \in Q} \big[\, \text{rel}(q) \cap \text{top}_k(q) \neq \varnothing \,\big] 154$$ 155 156**Symbols** 157 158| Symbol | Meaning here | Example | 159|---|---|---| 160| Q | the golden set of questions | 12 questions | 161| \|Q\| | how many questions | 12 | 162| q ∈ Q | each question in turn | | 163| rel(q) | the documents that truly answer q | {it-004} | 164| top_k(q) | the k documents search returned | 3 documents | 165| ∩, ≠ ∅ | "they share at least one document" | | 166| [ … ] | 1 if true, 0 if false | | 167| hit@k | the share of questions with a right answer in the top k | 0 to 1 | 168 169**In words:** the share of golden questions for which at least one right 170document appears among the top k results. (With one relevant document per 171question, as here, this equals recall@k.) 172 173**On the example:** 11 hits out of 12 = 0.92 with matching models; 3 out of 12 = 1740.25 with mixed models. 175 176**In Python:** 177 178```python 179def hit(rel_q, top_q): 180 # [rel(q) ∩ top_k(q) ≠ ∅] 181 return 1 if rel_q & top_q else 0 182hit({"it-004"}, {"it-002", "it-004", "fin-001"}) # → 1 183# matching models: 11 of the 12 questions 184hits = [1] * 11 + [0] 185# (1/|Q|) Σ over q in Q 186round(sum(hits) / len(hits), 2) # → 0.92 187# mixed models: 3 of the 12 188hits = [1] * 3 + [0] * 9 189round(sum(hits) / len(hits), 2) # → 0.25 190``` 191 192 193 194**Reading it:** each bar is recall@3 on the golden set. The first two bars 195use one model for both documents and queries, v1 then v2: both work. The 196third bar searches v1 documents with v2 queries: recall falls to 0.25 (3 of 197the 12 questions), barely above the dashed line at 0.15, which is what random 198ranking would score (3 of 20 documents shown). A model with a *different* 199number of dimensions would at least crash (you can't dot a 128-number vector 200with a 64-number one); a same-size model fails silently, which is worse. 201 202**In code:** `golden_recall` embeds the documents with one model and the 203questions with another and returns hit@k. `VectorIndex` is a flat index tied 204to one model version: `VectorIndex.search` embeds a text query, and 205`VectorIndex.search_vector` takes a vector that is already made. 206 207**Why it matters:** "we upgraded the embedding model and search got weird" 208is a classic incident. The cause is almost always documents and queries 209embedded by different models, typically because only new documents were 210re-embedded. 211 212## Migrating without a bad day: blue/green 213 214**Everyday picture:** building a new bridge next to the old one. Traffic 215keeps using the old bridge while the new one is built and inspected. When 216the new bridge passes inspection, traffic is switched over, and the old 217bridge stays standing for a while in case something turns up. 218 219```mermaid 220stateDiagram-v2 221 [*] --> Live_v1 222 Live_v1 --> DualWrite: start v2 (empty index, new docs go to both) 223 DualWrite --> Backfilled: backfill (re-embed old docs in batches) 224 Backfilled --> Compared: shadow compare (recall on golden set) 225 Compared --> Live_v2: cutover, only if v2 at least as good 226 Compared --> Live_v1: refused, v2 is worse 227 Live_v2 --> Live_v1: rollback (v1 index kept) 228 Live_v2 --> [*]: retire v1 after a soak period 229``` 230 231**Reading it:** follow the states from the top. Users are served by v1 the 232whole time until cutover. Dual-writing from the very first step means 233documents added mid-migration land in both indexes, so the new one never 234falls behind. The **shadow comparison** is the gate: the switch only happens 235if v2 is at least as good on your own golden set (`BlueGreenMigration` 236refuses to cut over without it). The old index is kept after cutover, so 237rollback is flipping one pointer, not a multi-hour rebuild. 238 239The "live alias" is a pointer: the search service asks for "the live 240index", and cutover or rollback just repoints it. Many vector databases 241support aliases for exactly this. 242 243**In code:** each arrow of the diagram is one method: `BlueGreenMigration.start`, 244`BlueGreenMigration.add` (the dual write), `BlueGreenMigration.backfill`, 245`BlueGreenMigration.shadow_compare`, `BlueGreenMigration.cutover` and 246`BlueGreenMigration.rollback`. 247 248**Tiny worked example: what will re-embedding cost?** 50 million documents of 249about 500 tokens each, an embedding throughput of 1 million tokens per second 250across your workers, and a price of \$0.02 per million tokens (all three are 251inputs you replace with your own numbers): 252 253$$ 254\text{hours} = \frac{n \cdot t}{r \cdot 3600}, \qquad 255\text{dollars} = \frac{n \cdot t}{10^6} \cdot p 256$$ 257 258**Symbols** 259 260| Symbol | Meaning here | Example | 261|---|---|---| 262| n | number of documents (or chunks) | 50,000,000 | 263| t | average tokens per document | 500 | 264| n · t | total tokens to embed | 25,000,000,000 | 265| r | throughput, tokens per second | 1,000,000 | 266| 3600 | seconds per hour | | 267| p | price per million tokens | \$0.02 | 268 269**In words:** total tokens divided by throughput gives the time; total 270tokens in millions times the price gives the cost. 271 272**On the example:** 25 × 10⁹ / (10⁶ × 3600) = **6.94 hours**; 25,000 × \$0.02 = 273**\$500**. 274 275**In Python:** 276 277```python 278# documents, tokens per document 279n, t = 50_000_000, 500 280# tokens per second, dollars per million tokens 281r, p = 1_000_000, 0.02 282# total tokens 283n * t # → 25000000000 284# hours 285round(n * t / (r * 3600), 2) # → 6.94 286# dollars 287round(n * t / 10**6 * p, 2) # → 500.0 288``` 289 290 291 292**Reading it:** the horizontal axis is corpus size (log scale); the left 293panel shows hours at three throughputs, the right panel shows dollars at 294three prices. Both grow in straight lines on these log axes: ten times the 295documents, ten times the time and money. The money is usually modest; the 296time, the rate limits and the double storage during the migration are what 297need planning. 298 299**In code:** `reembed_estimate` computes the tokens, hours and dollars for 300your own n, t, r and p. 301 302**Why it matters:** re-embedding is routine: new models, fine-tunes, 303chunking changes and bug fixes all require it. Versioned indexes, dual 304writes, a golden-set gate and a rollback path turn it from a risky event 305into a boring one. 306 307## Domain mismatch: fluent, but not in your jargon 308 309**Everyday picture:** a new hire who speaks perfect English but doesn't yet 310know that "AP" means accounts payable or that "T&E" means travel and 311expenses. General embedding models are trained mostly on web text; your 312acronyms and product names are new words to them. 313 314**Tiny worked example:** four real-sounding employee questions in company 315jargon: "ap aging report", "hotspot from a client site", "sso lockout", 316"t-and-e submission deadline". The general toy model puts the right 317document first for **1 of 4**. A version that has learned the four jargon 318words (`DomainTunedEmbedder`, standing in for a model fine-tuned on 319company pairs) gets **4 of 4**. 320 321 322 323**Reading it:** each pair of bars is one jargon question; the height is where 324the right document ranked (1 is best, shorter is better). The general model 325buries three of the four answers; the tuned model puts every one first. The 326one the general model gets right, "sso lockout", is saved by a word it does 327know ("lockout"). 328 329**In code:** `rank_of_first_relevant` finds where the right document lands 330for one question under one model: the height of each bar. 331 332```mermaid 333flowchart LR 334 L["Search & support logs"] --> F["Keep resolved sessions<br/>(the clicked doc answered it)"] 335 F --> N["Normalize and merge<br/>repeated questions"] 336 N --> G["Golden set: question → relevant docs<br/>(a few hundred is plenty)"] 337 G --> T["Test several models on YOUR set"] 338 T --> D{Good enough?} 339 D -->|no| FT["Fine-tune on domain pairs,<br/>add hybrid search"] 340 D -->|yes| S[Ship, and keep the set for regressions] 341 FT --> T 342``` 343 344**Reading it:** the evaluation set comes from real usage, not invention. 345Resolved sessions give you (question, answer document) pairs for free; 346`build_eval_set` keeps only resolved ones and merges repeats. Test 347candidate models on that set, and when none is good enough, fine-tune on 348domain pairs (`primer.ml.embeddings.contrastive`) and add keyword search, 349which matches jargon exactly (`primer.ml.embeddings.retrieval`). 350 351**Why it matters:** public leaderboards measure public data. The only 352number that predicts your system's quality is recall on your own 353questions. 354 355## Measure retrieval separately from generation 356 357**Everyday picture:** an open-book exam. A wrong answer has one of two 358causes: the right page wasn't in the book you brought (**retrieval 359failure**), or it was, and you misread it (**generation failure**). Studying 360harder fixes the second, not the first. 361 362**Tiny worked example:** three answered questions, checked by hand. 363 364| Retrieved (top 3) | Truly relevant | Answer used | Verdict | 365|---|---|---|---| 366| it-002, fin-006, fin-001 | it-004 | it-002 | retrieval failure: it-004 was never found | 367| hr-002, hr-001, hr-005 | hr-001 | hr-002 | generation failure: found but not used | 368| it-001, it-002 | it-001 | it-001 | correct | 369 370```mermaid 371flowchart TD 372 A[Wrong answer] --> R{Was a relevant document<br/>in the retrieved top k?} 373 R -->|no| RF[Retrieval failure:<br/>fix chunking, hybrid search,<br/>reranking, the embedding model] 374 R -->|yes| G{Did the answer use it?} 375 G -->|no| GF[Generation failure:<br/>fix the prompt, context order,<br/>fewer distracting chunks] 376 G -->|yes| OK[Check the grader:<br/>the answer may be fine] 377``` 378 379**Reading it:** one question splits every failure in two. If the right 380document wasn't retrieved, no prompt change can help, so work on retrieval. 381If it was retrieved and ignored, the problem is downstream, in the prompt or 382the context. Only the retrieval half can be measured without a model, which 383is why recall@k on a golden set is the first number to track. 384 385 386 387**Reading it:** each bar is the 12 golden questions, answered by a toy 388generator that always uses the top document, with k documents retrieved. 389Green is correct, red is a retrieval failure, orange is a generation 390failure. Retrieving more (larger k) turns some retrieval failures into 391generation failures: the right page is now in the book, but the reader still 392opened the wrong one. Watch the error-code question ("what does ERR-4012 393mean"): its answer ranks only fourth, so it's a retrieval failure until k = 5, 394and even then the top document is the wrong one. Dense vectors blur exact 395identifiers; keyword search would put that page first. 396 397**In code:** `triage` gives the verdict for one answered question, following 398the flowchart above; `triage_golden_set` runs every golden question through 399retrieval and a toy generator that answers from the top document. 400 401**Why it matters:** teams burn weeks tuning prompts for failures that were 402retrieval all along. Triage first, then fix the half that's broken. 403 404## In 20 seconds 405- Every embedding model has its own vector space; never mix documents and 406 queries from different models. Same-size models fail silently. 407- Migrate blue/green: new index, dual-write, backfill, compare on a golden 408 set, cut over only if better, keep the old index for rollback. 409- Re-embedding cost is n × tokens: estimate hours and dollars before you start. 410- General models miss company jargon; build an eval set from resolved logs, 411 and fine-tune or add keyword search when needed. 412- Triage RAG failures into retrieval vs. generation before fixing anything. 413 414## Self-test questions 415 416**Q: You need to switch embedding models on a 50-million-document index. Plan the migration.** 417Create a versioned index for the new model and dual-write new documents to 418both. Backfill by re-embedding existing documents in restartable batches 419(estimate: 50M × 500 tokens = 25B tokens; at 1M tokens/s that's about 7 hours). 420Compare recall on a golden set in shadow; cut over the live alias only if the 421new model is at least as good; keep the old index for fast rollback, then 422retire it. 423 424**Q: Why can't you compare vectors from two different embedding models?** 425Each model defines its own coordinate system. Even with the same number of 426dimensions, the same text lands in unrelated places, so similarity between 427the two spaces is meaningless. 428 429**Q: A general embedding model performs poorly on a company's internal documents. What helps?** 430Build a small eval set from real queries and the documents that resolved 431them, test several models on it, add hybrid (keyword + dense) search for 432exact terms, and fine-tune an embedding model on domain pairs with hard 433negatives if needed. 434 435**Q: Your RAG system gives confident wrong answers. What's your first diagnostic step?** 436Split retrieval from generation: for failing questions, check whether a 437relevant document was in the retrieved top k. If not, it's retrieval (fix 438search); if yes, it's generation (fix prompt and context). Track recall@k on 439a labeled set continuously. 440 441## The papers behind this lesson 442 443- **Thakur et al., *BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models* (2021)**: https://arxiv.org/abs/2104.08663. 444 Showed that retrieval models trained on one domain often lose badly on others, which is why you evaluate on your own data. 445- **Muennighoff et al., *MTEB: Massive Text Embedding Benchmark* (2022)**: https://arxiv.org/abs/2210.07316. 446 Compared embedding models across many tasks and found no single model wins everywhere. 447 448## Further reading 449- Es et al., *RAGAS: Automated Evaluation of Retrieval Augmented Generation* (2023): https://arxiv.org/abs/2309.15217 450- Martin Fowler, *BlueGreenDeployment*: https://martinfowler.com/bliki/BlueGreenDeployment.html 451- MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard 452""" 453 454from __future__ import annotations 455 456from dataclasses import dataclass, field 457 458import numpy as np 459 460from primer._show import banner, say, table, takeaway 461from primer.common.corpus import DOCS, LABELED_QUERIES, Doc 462from primer.common.embedder import ConceptEmbedder, _unit_vector 463from primer.common.text import tokenize 464 465# Model versions. Each `name` is a different model, so each defines its own 466# vector space, even when two have the same number of dimensions. 467CURRENT_MODEL = ConceptEmbedder() # the repo's standard toy model, "concept-v1", 64 dims 468SAME_SIZE_NEW_MODEL = ConceptEmbedder(name="concept-v2") # also 64 dims: mixing it with v1 fails *silently* 469BETTER_MODEL = ConceptEmbedder(dim=128, name="concept-v2-large") 470WORSE_MODEL = ConceptEmbedder(dim=8, name="concept-v2-mini") 471 472 473def doc_text(d: Doc) -> str: 474 return f"{d.title}. {d.text}" 475 476 477# --------------------------------------------------------------------------- 478# 1. A minimal versioned vector index, and golden-set recall 479# --------------------------------------------------------------------------- 480 481 482@dataclass 483class VectorIndex: 484 """Exact (flat) index for one model version. Vectors are only comparable within it.""" 485 486 model: ConceptEmbedder 487 ids: list[str] = field(default_factory=list) 488 vectors: list[np.ndarray] = field(default_factory=list) 489 490 @property 491 def version(self) -> str: 492 return self.model.name 493 494 def add(self, docs: list[Doc]) -> None: 495 for d, v in zip(docs, self.model.encode([doc_text(d) for d in docs])): 496 self.ids.append(d.id) 497 self.vectors.append(v) 498 499 def search_vector(self, q: np.ndarray, k: int) -> list[str]: 500 scores = np.stack(self.vectors) @ q 501 return [self.ids[i] for i in np.argsort(-scores, kind="stable")[:k]] 502 503 def search(self, query: str, k: int) -> list[str]: 504 return self.search_vector(self.model.encode(query), k) 505 506 507def golden_recall(query_model: ConceptEmbedder, doc_model: ConceptEmbedder, k: int, queries=LABELED_QUERIES) -> float: 508 """Share of golden queries with at least one relevant doc in the top k. 509 510 Documents are embedded with `doc_model` and queries with `query_model`. 511 Passing two different models reproduces the classic migration bug: 512 querying an old index with a new model. 513 """ 514 idx = VectorIndex(doc_model) 515 idx.add(DOCS) 516 hits = [bool(rel & set(idx.search_vector(query_model.encode(q), k))) for q, rel in queries] 517 return float(np.mean(hits)) 518 519 520# --------------------------------------------------------------------------- 521# 2. Blue/green migration to a new embedding model 522# --------------------------------------------------------------------------- 523 524 525class BlueGreenMigration: 526 """Move search from one embedding model to another without a bad day. 527 528 Steps: start (create the new index and begin dual-writing), backfill 529 (re-embed every existing document), shadow_compare (measure both on the 530 golden set), cutover (flip the live alias only if the new one is at least 531 as good), rollback (flip it back; the old index is kept until retired). 532 """ 533 534 def __init__(self, current: ConceptEmbedder, docs: list[Doc] = DOCS): 535 self.docs = list(docs) 536 self.indexes: dict[str, VectorIndex] = {current.name: VectorIndex(current)} 537 self.indexes[current.name].add(self.docs) 538 self.live_version = current.name 539 self.previous_version: str | None = None 540 self.candidate: str | None = None 541 self.comparison: dict[str, float] = {} 542 self.log: list[str] = [f"live index {current.name} serving {len(self.docs)} docs"] 543 544 def start(self, new: ConceptEmbedder) -> None: 545 self.indexes[new.name] = VectorIndex(new) 546 self.candidate = new.name 547 self.comparison = {} 548 self.log.append(f"created empty index {new.name}; dual-writing new documents to both") 549 550 def add(self, doc: Doc) -> None: 551 """New documents go to every index, so the candidate never falls behind while backfilling.""" 552 self.docs.append(doc) 553 for idx in self.indexes.values(): 554 idx.add([doc]) 555 self.log.append(f"dual-wrote {doc.id} to {sorted(self.indexes)}") 556 557 def backfill(self, batch_size: int = 8) -> None: 558 idx = self.indexes[self.candidate] 559 missing = [d for d in self.docs if d.id not in set(idx.ids)] 560 for start in range(0, len(missing), batch_size): # batches: restartable, rate-limit friendly 561 idx.add(missing[start : start + batch_size]) 562 self.log.append(f"backfilled {len(missing)} docs into {self.candidate}") 563 564 def shadow_compare(self, k: int = 3) -> dict[str, float]: 565 for name in (self.live_version, self.candidate): 566 m = self.indexes[name].model 567 self.comparison[name] = golden_recall(m, m, k) 568 self.log.append("shadow recall@%d: %s" % (k, ", ".join(f"{n} {r:.3f}" for n, r in self.comparison.items()))) 569 return self.comparison 570 571 def cutover(self, min_gain: float = 0.0) -> bool: 572 if self.candidate not in self.comparison: 573 self.log.append("cutover refused: no shadow comparison yet") 574 return False 575 old, new = self.comparison[self.live_version], self.comparison[self.candidate] 576 if new < old + min_gain: 577 self.log.append(f"cutover refused: {self.candidate} recall {new:.3f} < live {old:.3f}") 578 return False 579 self.previous_version, self.live_version = self.live_version, self.candidate 580 self.log.append(f"cutover: live alias -> {self.live_version} (old index kept for rollback)") 581 return True 582 583 def rollback(self) -> None: 584 if self.previous_version: 585 self.live_version, self.previous_version = self.previous_version, self.live_version 586 self.log.append(f"rollback: live alias -> {self.live_version}") 587 588 589def reembed_estimate(n_docs: int, avg_tokens: int, tokens_per_second: float, dollars_per_million_tokens: float) -> dict[str, float]: 590 """Back-of-envelope time and money to re-embed a corpus. Prices vary; pass your own.""" 591 tokens = n_docs * avg_tokens 592 return {"tokens": tokens, "hours": tokens / tokens_per_second / 3600, "dollars": tokens / 1e6 * dollars_per_million_tokens} 593 594 595# --------------------------------------------------------------------------- 596# 3. Domain mismatch 597# --------------------------------------------------------------------------- 598 599# Jargon a general model has never seen, and the concept each word means here. 600JARGON: dict[str, str] = {"ap": "invoice", "hotspot": "remote_access", "sso": "credential", "t-and-e": "money_back"} 601 602JARGON_QUERIES: list[tuple[str, set[str]]] = [ 603 ("ap aging report", {"fin-005"}), # accounts payable: vendor invoices 604 ("hotspot from a client site", {"it-003"}), # remote access 605 ("sso lockout", {"it-001"}), # single sign-on: your login 606 ("t-and-e submission deadline", {"fin-004"}), # travel and expenses 607] 608 609 610class DomainTunedEmbedder(ConceptEmbedder): 611 """Stand-in for a model fine-tuned on company pairs: it has learned the jargon. 612 613 In real life this comes from contrastive fine-tuning on (query, document) 614 pairs from your own logs (`primer.ml.embeddings.contrastive`). Here we 615 simply teach the toy embedder what four jargon words mean. 616 """ 617 618 def __init__(self, jargon: dict[str, str] = JARGON, **kw): 619 super().__init__(**{"name": "concept-v1-tuned", **kw}) 620 self.jargon = jargon 621 622 def _token_vector(self, tok: str) -> np.ndarray: 623 concept = self.jargon.get(tok) 624 if concept is None: 625 return super()._token_vector(tok) 626 return _unit_vector(f"{self.name}/concept/{concept}", self.dim) + self.word_weight * _unit_vector(f"{self.name}/word/{tok}", self.dim) 627 628 629def build_eval_set(log: list[dict]) -> list[tuple[str, set[str]]]: 630 """Turn a search/support log into (query, relevant doc ids) cases. 631 632 Keep only sessions where the user's problem was resolved (the clicked 633 document really answered it); merge the same question asked twice. 634 """ 635 cases: dict[str, set[str]] = {} 636 for row in log: 637 if not row["resolved"]: 638 continue 639 key = " ".join(tokenize(row["query"], drop_stopwords=False)) # "Travel policy?" == "travel policy" 640 cases.setdefault(key, set()).add(row["clicked"]) 641 return list(cases.items()) 642 643 644# --------------------------------------------------------------------------- 645# 4. Measure retrieval separately from generation 646# --------------------------------------------------------------------------- 647 648 649def triage(retrieved: list[str], relevant: set[str], answer_used: str) -> str: 650 """Which half of a RAG system failed? 651 652 If no relevant document was retrieved, no prompt can fix the answer: it's 653 a retrieval failure. If one was retrieved but the answer relied on a 654 different document, the generator misused good context. 655 """ 656 if not relevant & set(retrieved): 657 return "retrieval failure" 658 if answer_used not in relevant: 659 return "generation failure" 660 return "correct" 661 662 663def triage_golden_set(model: ConceptEmbedder, k: int = 3, queries=LABELED_QUERIES) -> list[tuple[str, str]]: 664 """Run every golden query through retrieval and a toy generator that answers from the top document.""" 665 idx = VectorIndex(model) 666 idx.add(DOCS) 667 out = [] 668 for q, rel in queries: 669 top = idx.search(q, k) 670 out.append((q, triage(top, rel, answer_used=top[0]))) 671 return out 672 673 674 675def rank_of_first_relevant(model: ConceptEmbedder, query: str, relevant: set[str]) -> int: 676 """1-based position of the first relevant document in the full ranking.""" 677 idx = VectorIndex(model) 678 idx.add(DOCS) 679 ranking = idx.search(query, len(DOCS)) 680 return next(i + 1 for i, d in enumerate(ranking) if d in relevant) 681 682 683# --------------------------------------------------------------------------- 684# 5. Figures 685# --------------------------------------------------------------------------- 686 687 688def figures() -> dict: 689 """Plots computed from this module's own functions. Keys match the docstring's image names.""" 690 import matplotlib 691 692 matplotlib.use("Agg") 693 import matplotlib.pyplot as plt 694 695 figs = {} 696 697 # cross-model 698 labels = ["v1 docs,\nv1 queries", "v2 docs,\nv2 queries", "v1 docs,\nv2 queries"] 699 vals = [golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)] 700 fig, ax = plt.subplots(figsize=(6, 4)) 701 bars = ax.bar(labels, vals, color=["C0", "C2", "C3"]) 702 ax.bar_label(bars, fmt="%.2f") 703 ax.axhline(3 / len(DOCS), ls="--", color="0.5", label="random ranking (3 of 20)") 704 ax.set(ylabel="recall@3 on the golden set", ylim=(0, 1.1), title="Mixing two models' vectors fails silently") 705 ax.legend() 706 figs["cross_model"] = fig 707 708 # reembed 709 ns = np.logspace(5, 9, 30) 710 fig, (a1, a2) = plt.subplots(1, 2, figsize=(11, 4)) 711 for r in (1e5, 1e6, 1e7): 712 a1.plot(ns, [reembed_estimate(int(n), 500, r, 0.02)["hours"] for n in ns], label=f"{r:,.0f} tokens/s") 713 for p in (0.01, 0.02, 0.13): 714 a2.plot(ns, [reembed_estimate(int(n), 500, 1e6, p)["dollars"] for n in ns], label=f"${p} per million tokens") 715 for ax, ylab in ((a1, "hours"), (a2, "dollars")): 716 ax.set_xscale("log") 717 ax.set_yscale("log") 718 ax.axvline(5e7, ls=":", color="0.5") 719 ax.set(xlabel="documents (500 tokens each, log scale)", ylabel=f"{ylab} (log scale)") 720 ax.legend(fontsize=8) 721 a1.set_title("Time to re-embed") 722 a2.set_title("Cost to re-embed") 723 figs["reembed"] = fig 724 725 # jargon 726 tuned = DomainTunedEmbedder() 727 qs = [q for q, _ in JARGON_QUERIES] 728 x = np.arange(len(qs)) 729 fig, ax = plt.subplots(figsize=(8, 4)) 730 ax.bar(x - 0.2, [rank_of_first_relevant(CURRENT_MODEL, q, r) for q, r in JARGON_QUERIES], 0.4, label="general model") 731 ax.bar(x + 0.2, [rank_of_first_relevant(tuned, q, r) for q, r in JARGON_QUERIES], 0.4, label="tuned on the jargon") 732 ax.set_xticks(x, qs, fontsize=8) 733 ax.set(ylabel="rank of the right document (1 is best)", title="Company jargon: general vs. tuned model") 734 ax.legend() 735 figs["jargon"] = fig 736 737 # triage 738 ks = [1, 3, 5] 739 colours = {"correct": "C2", "generation failure": "C1", "retrieval failure": "C3"} 740 fig, ax = plt.subplots(figsize=(6, 4)) 741 bottoms = np.zeros(len(ks)) 742 counts = {v: [] for v in colours} 743 for k in ks: 744 verdicts = [v for _, v in triage_golden_set(CURRENT_MODEL, k)] 745 for v in colours: 746 counts[v].append(verdicts.count(v)) 747 for v, c in colours.items(): 748 ax.bar([str(k) for k in ks], counts[v], bottom=bottoms, color=c, label=v) 749 bottoms += np.array(counts[v]) 750 ax.set(xlabel="documents retrieved (k)", ylabel="golden questions", title="Which half failed? (generator uses the top document)") 751 ax.legend(fontsize=8) 752 figs["triage"] = fig 753 754 for f in figs.values(): 755 f.tight_layout() 756 return figs 757 758 759# --------------------------------------------------------------------------- 760# 6. Narrated walkthrough 761# --------------------------------------------------------------------------- 762 763 764def demo() -> None: 765 banner("1. Every model has its own space") 766 a, b = ConceptEmbedder(name="concept-v1").encode("vpn"), ConceptEmbedder(name="concept-v2").encode("vpn") 767 say(f"Cosine between 'vpn' from v1 and 'vpn' from v2 (both 64 dims): {float(a @ b):.3f}.") 768 table( 769 ["documents", "queries", "recall@3"], 770 [ 771 ("v1", "v1", golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3)), 772 ("v2", "v2", golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3)), 773 ("v1", "v2 (the bug)", golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)), 774 ], 775 floatfmt=".3f", 776 ) 777 takeaway("Same number of dimensions, different space. Mixing them doesn't crash; it just quietly stops working.") 778 779 banner("2. A blue/green migration, step by step") 780 for candidate in (BETTER_MODEL, WORSE_MODEL): 781 m = BlueGreenMigration(CURRENT_MODEL) 782 m.start(candidate) 783 m.add(Doc("it-999", "Wi-Fi guest network", "The guest Wi-Fi password rotates every Monday.", "IT", "2026-09-01")) 784 m.backfill() 785 m.shadow_compare(k=3) 786 m.cutover() 787 print(f" candidate {candidate.name}:") 788 for line in m.log: 789 print(" -", line) 790 print() 791 est = reembed_estimate(50_000_000, 500, 1_000_000, 0.02) 792 say(f"Re-embedding 50M docs × 500 tokens = {est['tokens']:,} tokens: {est['hours']:.2f} hours at 1M tokens/s, ${est['dollars']:,.0f} at $0.02 per million.") 793 794 banner("3. Domain jargon") 795 tuned = DomainTunedEmbedder() 796 table( 797 ["jargon question", "rank (general)", "rank (tuned)"], 798 [(q, rank_of_first_relevant(CURRENT_MODEL, q, r), rank_of_first_relevant(tuned, q, r)) for q, r in JARGON_QUERIES], 799 ) 800 log = [ 801 {"query": "VPN error 4012?", "clicked": "it-004", "resolved": True}, 802 {"query": "vpn error 4012", "clicked": "it-003", "resolved": True}, 803 {"query": "printer", "clicked": "it-009", "resolved": False}, 804 ] 805 say(f"Eval set built from a 3-line log: {build_eval_set(log)}") 806 807 banner("4. Triage: retrieval failure or generation failure?") 808 table(["golden question", "verdict (k = 3)"], triage_golden_set(CURRENT_MODEL, 3)) 809 takeaway("If the right document wasn't retrieved, no prompt can fix it. Measure retrieval on its own first.") 810 811 812if __name__ == "__main__": 813 demo()
483@dataclass 484class VectorIndex: 485 """Exact (flat) index for one model version. Vectors are only comparable within it.""" 486 487 model: ConceptEmbedder 488 ids: list[str] = field(default_factory=list) 489 vectors: list[np.ndarray] = field(default_factory=list) 490 491 @property 492 def version(self) -> str: 493 return self.model.name 494 495 def add(self, docs: list[Doc]) -> None: 496 for d, v in zip(docs, self.model.encode([doc_text(d) for d in docs])): 497 self.ids.append(d.id) 498 self.vectors.append(v) 499 500 def search_vector(self, q: np.ndarray, k: int) -> list[str]: 501 scores = np.stack(self.vectors) @ q 502 return [self.ids[i] for i in np.argsort(-scores, kind="stable")[:k]] 503 504 def search(self, query: str, k: int) -> list[str]: 505 return self.search_vector(self.model.encode(query), k)
Exact (flat) index for one model version. Vectors are only comparable within it.
508def golden_recall(query_model: ConceptEmbedder, doc_model: ConceptEmbedder, k: int, queries=LABELED_QUERIES) -> float: 509 """Share of golden queries with at least one relevant doc in the top k. 510 511 Documents are embedded with `doc_model` and queries with `query_model`. 512 Passing two different models reproduces the classic migration bug: 513 querying an old index with a new model. 514 """ 515 idx = VectorIndex(doc_model) 516 idx.add(DOCS) 517 hits = [bool(rel & set(idx.search_vector(query_model.encode(q), k))) for q, rel in queries] 518 return float(np.mean(hits))
Share of golden queries with at least one relevant doc in the top k.
Documents are embedded with doc_model and queries with query_model.
Passing two different models reproduces the classic migration bug:
querying an old index with a new model.
526class BlueGreenMigration: 527 """Move search from one embedding model to another without a bad day. 528 529 Steps: start (create the new index and begin dual-writing), backfill 530 (re-embed every existing document), shadow_compare (measure both on the 531 golden set), cutover (flip the live alias only if the new one is at least 532 as good), rollback (flip it back; the old index is kept until retired). 533 """ 534 535 def __init__(self, current: ConceptEmbedder, docs: list[Doc] = DOCS): 536 self.docs = list(docs) 537 self.indexes: dict[str, VectorIndex] = {current.name: VectorIndex(current)} 538 self.indexes[current.name].add(self.docs) 539 self.live_version = current.name 540 self.previous_version: str | None = None 541 self.candidate: str | None = None 542 self.comparison: dict[str, float] = {} 543 self.log: list[str] = [f"live index {current.name} serving {len(self.docs)} docs"] 544 545 def start(self, new: ConceptEmbedder) -> None: 546 self.indexes[new.name] = VectorIndex(new) 547 self.candidate = new.name 548 self.comparison = {} 549 self.log.append(f"created empty index {new.name}; dual-writing new documents to both") 550 551 def add(self, doc: Doc) -> None: 552 """New documents go to every index, so the candidate never falls behind while backfilling.""" 553 self.docs.append(doc) 554 for idx in self.indexes.values(): 555 idx.add([doc]) 556 self.log.append(f"dual-wrote {doc.id} to {sorted(self.indexes)}") 557 558 def backfill(self, batch_size: int = 8) -> None: 559 idx = self.indexes[self.candidate] 560 missing = [d for d in self.docs if d.id not in set(idx.ids)] 561 for start in range(0, len(missing), batch_size): # batches: restartable, rate-limit friendly 562 idx.add(missing[start : start + batch_size]) 563 self.log.append(f"backfilled {len(missing)} docs into {self.candidate}") 564 565 def shadow_compare(self, k: int = 3) -> dict[str, float]: 566 for name in (self.live_version, self.candidate): 567 m = self.indexes[name].model 568 self.comparison[name] = golden_recall(m, m, k) 569 self.log.append("shadow recall@%d: %s" % (k, ", ".join(f"{n} {r:.3f}" for n, r in self.comparison.items()))) 570 return self.comparison 571 572 def cutover(self, min_gain: float = 0.0) -> bool: 573 if self.candidate not in self.comparison: 574 self.log.append("cutover refused: no shadow comparison yet") 575 return False 576 old, new = self.comparison[self.live_version], self.comparison[self.candidate] 577 if new < old + min_gain: 578 self.log.append(f"cutover refused: {self.candidate} recall {new:.3f} < live {old:.3f}") 579 return False 580 self.previous_version, self.live_version = self.live_version, self.candidate 581 self.log.append(f"cutover: live alias -> {self.live_version} (old index kept for rollback)") 582 return True 583 584 def rollback(self) -> None: 585 if self.previous_version: 586 self.live_version, self.previous_version = self.previous_version, self.live_version 587 self.log.append(f"rollback: live alias -> {self.live_version}")
Move search from one embedding model to another without a bad day.
Steps: start (create the new index and begin dual-writing), backfill (re-embed every existing document), shadow_compare (measure both on the golden set), cutover (flip the live alias only if the new one is at least as good), rollback (flip it back; the old index is kept until retired).
535 def __init__(self, current: ConceptEmbedder, docs: list[Doc] = DOCS): 536 self.docs = list(docs) 537 self.indexes: dict[str, VectorIndex] = {current.name: VectorIndex(current)} 538 self.indexes[current.name].add(self.docs) 539 self.live_version = current.name 540 self.previous_version: str | None = None 541 self.candidate: str | None = None 542 self.comparison: dict[str, float] = {} 543 self.log: list[str] = [f"live index {current.name} serving {len(self.docs)} docs"]
551 def add(self, doc: Doc) -> None: 552 """New documents go to every index, so the candidate never falls behind while backfilling.""" 553 self.docs.append(doc) 554 for idx in self.indexes.values(): 555 idx.add([doc]) 556 self.log.append(f"dual-wrote {doc.id} to {sorted(self.indexes)}")
New documents go to every index, so the candidate never falls behind while backfilling.
558 def backfill(self, batch_size: int = 8) -> None: 559 idx = self.indexes[self.candidate] 560 missing = [d for d in self.docs if d.id not in set(idx.ids)] 561 for start in range(0, len(missing), batch_size): # batches: restartable, rate-limit friendly 562 idx.add(missing[start : start + batch_size]) 563 self.log.append(f"backfilled {len(missing)} docs into {self.candidate}")
565 def shadow_compare(self, k: int = 3) -> dict[str, float]: 566 for name in (self.live_version, self.candidate): 567 m = self.indexes[name].model 568 self.comparison[name] = golden_recall(m, m, k) 569 self.log.append("shadow recall@%d: %s" % (k, ", ".join(f"{n} {r:.3f}" for n, r in self.comparison.items()))) 570 return self.comparison
572 def cutover(self, min_gain: float = 0.0) -> bool: 573 if self.candidate not in self.comparison: 574 self.log.append("cutover refused: no shadow comparison yet") 575 return False 576 old, new = self.comparison[self.live_version], self.comparison[self.candidate] 577 if new < old + min_gain: 578 self.log.append(f"cutover refused: {self.candidate} recall {new:.3f} < live {old:.3f}") 579 return False 580 self.previous_version, self.live_version = self.live_version, self.candidate 581 self.log.append(f"cutover: live alias -> {self.live_version} (old index kept for rollback)") 582 return True
590def reembed_estimate(n_docs: int, avg_tokens: int, tokens_per_second: float, dollars_per_million_tokens: float) -> dict[str, float]: 591 """Back-of-envelope time and money to re-embed a corpus. Prices vary; pass your own.""" 592 tokens = n_docs * avg_tokens 593 return {"tokens": tokens, "hours": tokens / tokens_per_second / 3600, "dollars": tokens / 1e6 * dollars_per_million_tokens}
Back-of-envelope time and money to re-embed a corpus. Prices vary; pass your own.
611class DomainTunedEmbedder(ConceptEmbedder): 612 """Stand-in for a model fine-tuned on company pairs: it has learned the jargon. 613 614 In real life this comes from contrastive fine-tuning on (query, document) 615 pairs from your own logs (`primer.ml.embeddings.contrastive`). Here we 616 simply teach the toy embedder what four jargon words mean. 617 """ 618 619 def __init__(self, jargon: dict[str, str] = JARGON, **kw): 620 super().__init__(**{"name": "concept-v1-tuned", **kw}) 621 self.jargon = jargon 622 623 def _token_vector(self, tok: str) -> np.ndarray: 624 concept = self.jargon.get(tok) 625 if concept is None: 626 return super()._token_vector(tok) 627 return _unit_vector(f"{self.name}/concept/{concept}", self.dim) + self.word_weight * _unit_vector(f"{self.name}/word/{tok}", self.dim)
Stand-in for a model fine-tuned on company pairs: it has learned the jargon.
In real life this comes from contrastive fine-tuning on (query, document)
pairs from your own logs (primer.ml.embeddings.contrastive). Here we
simply teach the toy embedder what four jargon words mean.
Inherited Members
630def build_eval_set(log: list[dict]) -> list[tuple[str, set[str]]]: 631 """Turn a search/support log into (query, relevant doc ids) cases. 632 633 Keep only sessions where the user's problem was resolved (the clicked 634 document really answered it); merge the same question asked twice. 635 """ 636 cases: dict[str, set[str]] = {} 637 for row in log: 638 if not row["resolved"]: 639 continue 640 key = " ".join(tokenize(row["query"], drop_stopwords=False)) # "Travel policy?" == "travel policy" 641 cases.setdefault(key, set()).add(row["clicked"]) 642 return list(cases.items())
Turn a search/support log into (query, relevant doc ids) cases.
Keep only sessions where the user's problem was resolved (the clicked document really answered it); merge the same question asked twice.
650def triage(retrieved: list[str], relevant: set[str], answer_used: str) -> str: 651 """Which half of a RAG system failed? 652 653 If no relevant document was retrieved, no prompt can fix the answer: it's 654 a retrieval failure. If one was retrieved but the answer relied on a 655 different document, the generator misused good context. 656 """ 657 if not relevant & set(retrieved): 658 return "retrieval failure" 659 if answer_used not in relevant: 660 return "generation failure" 661 return "correct"
Which half of a RAG system failed?
If no relevant document was retrieved, no prompt can fix the answer: it's a retrieval failure. If one was retrieved but the answer relied on a different document, the generator misused good context.
664def triage_golden_set(model: ConceptEmbedder, k: int = 3, queries=LABELED_QUERIES) -> list[tuple[str, str]]: 665 """Run every golden query through retrieval and a toy generator that answers from the top document.""" 666 idx = VectorIndex(model) 667 idx.add(DOCS) 668 out = [] 669 for q, rel in queries: 670 top = idx.search(q, k) 671 out.append((q, triage(top, rel, answer_used=top[0]))) 672 return out
Run every golden query through retrieval and a toy generator that answers from the top document.
676def rank_of_first_relevant(model: ConceptEmbedder, query: str, relevant: set[str]) -> int: 677 """1-based position of the first relevant document in the full ranking.""" 678 idx = VectorIndex(model) 679 idx.add(DOCS) 680 ranking = idx.search(query, len(DOCS)) 681 return next(i + 1 for i, d in enumerate(ranking) if d in relevant)
1-based position of the first relevant document in the full ranking.
689def figures() -> dict: 690 """Plots computed from this module's own functions. Keys match the docstring's image names.""" 691 import matplotlib 692 693 matplotlib.use("Agg") 694 import matplotlib.pyplot as plt 695 696 figs = {} 697 698 # cross-model 699 labels = ["v1 docs,\nv1 queries", "v2 docs,\nv2 queries", "v1 docs,\nv2 queries"] 700 vals = [golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)] 701 fig, ax = plt.subplots(figsize=(6, 4)) 702 bars = ax.bar(labels, vals, color=["C0", "C2", "C3"]) 703 ax.bar_label(bars, fmt="%.2f") 704 ax.axhline(3 / len(DOCS), ls="--", color="0.5", label="random ranking (3 of 20)") 705 ax.set(ylabel="recall@3 on the golden set", ylim=(0, 1.1), title="Mixing two models' vectors fails silently") 706 ax.legend() 707 figs["cross_model"] = fig 708 709 # reembed 710 ns = np.logspace(5, 9, 30) 711 fig, (a1, a2) = plt.subplots(1, 2, figsize=(11, 4)) 712 for r in (1e5, 1e6, 1e7): 713 a1.plot(ns, [reembed_estimate(int(n), 500, r, 0.02)["hours"] for n in ns], label=f"{r:,.0f} tokens/s") 714 for p in (0.01, 0.02, 0.13): 715 a2.plot(ns, [reembed_estimate(int(n), 500, 1e6, p)["dollars"] for n in ns], label=f"${p} per million tokens") 716 for ax, ylab in ((a1, "hours"), (a2, "dollars")): 717 ax.set_xscale("log") 718 ax.set_yscale("log") 719 ax.axvline(5e7, ls=":", color="0.5") 720 ax.set(xlabel="documents (500 tokens each, log scale)", ylabel=f"{ylab} (log scale)") 721 ax.legend(fontsize=8) 722 a1.set_title("Time to re-embed") 723 a2.set_title("Cost to re-embed") 724 figs["reembed"] = fig 725 726 # jargon 727 tuned = DomainTunedEmbedder() 728 qs = [q for q, _ in JARGON_QUERIES] 729 x = np.arange(len(qs)) 730 fig, ax = plt.subplots(figsize=(8, 4)) 731 ax.bar(x - 0.2, [rank_of_first_relevant(CURRENT_MODEL, q, r) for q, r in JARGON_QUERIES], 0.4, label="general model") 732 ax.bar(x + 0.2, [rank_of_first_relevant(tuned, q, r) for q, r in JARGON_QUERIES], 0.4, label="tuned on the jargon") 733 ax.set_xticks(x, qs, fontsize=8) 734 ax.set(ylabel="rank of the right document (1 is best)", title="Company jargon: general vs. tuned model") 735 ax.legend() 736 figs["jargon"] = fig 737 738 # triage 739 ks = [1, 3, 5] 740 colours = {"correct": "C2", "generation failure": "C1", "retrieval failure": "C3"} 741 fig, ax = plt.subplots(figsize=(6, 4)) 742 bottoms = np.zeros(len(ks)) 743 counts = {v: [] for v in colours} 744 for k in ks: 745 verdicts = [v for _, v in triage_golden_set(CURRENT_MODEL, k)] 746 for v in colours: 747 counts[v].append(verdicts.count(v)) 748 for v, c in colours.items(): 749 ax.bar([str(k) for k in ks], counts[v], bottom=bottoms, color=c, label=v) 750 bottoms += np.array(counts[v]) 751 ax.set(xlabel="documents retrieved (k)", ylabel="golden questions", title="Which half failed? (generator uses the top document)") 752 ax.legend(fontsize=8) 753 figs["triage"] = fig 754 755 for f in figs.values(): 756 f.tight_layout() 757 return figs
Plots computed from this module's own functions. Keys match the docstring's image names.
765def demo() -> None: 766 banner("1. Every model has its own space") 767 a, b = ConceptEmbedder(name="concept-v1").encode("vpn"), ConceptEmbedder(name="concept-v2").encode("vpn") 768 say(f"Cosine between 'vpn' from v1 and 'vpn' from v2 (both 64 dims): {float(a @ b):.3f}.") 769 table( 770 ["documents", "queries", "recall@3"], 771 [ 772 ("v1", "v1", golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3)), 773 ("v2", "v2", golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3)), 774 ("v1", "v2 (the bug)", golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)), 775 ], 776 floatfmt=".3f", 777 ) 778 takeaway("Same number of dimensions, different space. Mixing them doesn't crash; it just quietly stops working.") 779 780 banner("2. A blue/green migration, step by step") 781 for candidate in (BETTER_MODEL, WORSE_MODEL): 782 m = BlueGreenMigration(CURRENT_MODEL) 783 m.start(candidate) 784 m.add(Doc("it-999", "Wi-Fi guest network", "The guest Wi-Fi password rotates every Monday.", "IT", "2026-09-01")) 785 m.backfill() 786 m.shadow_compare(k=3) 787 m.cutover() 788 print(f" candidate {candidate.name}:") 789 for line in m.log: 790 print(" -", line) 791 print() 792 est = reembed_estimate(50_000_000, 500, 1_000_000, 0.02) 793 say(f"Re-embedding 50M docs × 500 tokens = {est['tokens']:,} tokens: {est['hours']:.2f} hours at 1M tokens/s, ${est['dollars']:,.0f} at $0.02 per million.") 794 795 banner("3. Domain jargon") 796 tuned = DomainTunedEmbedder() 797 table( 798 ["jargon question", "rank (general)", "rank (tuned)"], 799 [(q, rank_of_first_relevant(CURRENT_MODEL, q, r), rank_of_first_relevant(tuned, q, r)) for q, r in JARGON_QUERIES], 800 ) 801 log = [ 802 {"query": "VPN error 4012?", "clicked": "it-004", "resolved": True}, 803 {"query": "vpn error 4012", "clicked": "it-003", "resolved": True}, 804 {"query": "printer", "clicked": "it-009", "resolved": False}, 805 ] 806 say(f"Eval set built from a 3-line log: {build_eval_set(log)}") 807 808 banner("4. Triage: retrieval failure or generation failure?") 809 table(["golden question", "verdict (k = 3)"], triage_golden_set(CURRENT_MODEL, 3)) 810 takeaway("If the right document wasn't retrieved, no prompt can fix it. Measure retrieval on its own first.")