primer.ml.embeddings.operations

Operating embeddings: model changes, jargon, and finding what broke

Run: python -m primer.ml.embeddings.operations

New to vectors or Σ? primer.notation builds them from zero. This lesson builds on primer.ml.embeddings.similarity and primer.ml.embeddings.contrastive.

Level 1: The practitioner's guide

In one sentence. Operating embeddings means keeping every query on the same map as the documents it searches, changing that map without a bad day, checking that the map speaks your domain's language, and knowing which half of a retrieval-augmented system failed when an answer is wrong.

When you need it. From the day an embedding index serves real traffic. Every embedding model is its own coordinate system: in this lesson, the word "vpn" embedded by two versions of the same toy model, both 64 dimensions, has a cosine similarity of about 0.05, as if they were unrelated words. Search documents embedded by one model with queries embedded by another and recall@3 on this lesson's golden set falls from 0.92 to 0.25, barely above the 0.15 that random ranking would score, and nothing crashes. You will change models: new releases, a fine-tune, a chunking change and a bug fix all require re-embedding everything. You will meet jargon the model doesn't know: this lesson's general model puts the right document first for one of four company-jargon questions. And you will get a confident wrong answer and need to know whether retrieval or generation caused it. The tell: "we upgraded the embedding model and search got weird", or weeks spent tuning prompts for answers whose right page was never retrieved.

Your options. Two decisions recur: how to change the model, and what to do when it doesn't know your words.

Option What it does What it gives you What it costs Where it lives
Re-embed only new documents Old documents keep their old vectors; new ones get the new model Nothing: it is the classic incident. Queries and documents stop sharing a space and recall collapses silently The cheapest to run and the most expensive to discover A pipeline that forgot the backfill
Re-embed in place Stops writes, re-embeds every document into the same index, resumes A correct index at the end Downtime or a window of mixed results, and no rollback except re-embedding again Your ingestion job
Blue/green migration Builds a second index alongside, dual-writes new documents to both, backfills, compares on a golden set, then repoints a live alias No downtime, a gate that refuses a worse model, rollback by flipping one pointer Double storage during the migration, a golden set, the backfill time Your code plus index aliases (Elasticsearch, Qdrant)
Keep the general model, add keyword search Fuses BM25 with dense search so exact jargon matches by spelling Identifiers and acronyms found without training Two searches per query Any search engine (primer.ml.embeddings.retrieval)
Fine-tune on domain pairs Trains the embedding model on your own (question, document) pairs Four of four jargon questions right in this lesson, against one of four Labeled pairs, a training run, then a full re-embed primer.ml.embeddings.contrastive, followed by a migration
Measure and triage Keeps a golden set of questions with known answer passages and checks it on every change Silent failures made visible, each split into retrieval or generation A few hundred labeled questions, taken from resolved logs Your evaluation harness; Ragas

How to choose.

  • Changing the model on a live index: blue/green, every time. Version the index by model, dual-write from the first minute so nothing falls behind, backfill in restartable batches, and cut over only when the new index is at least as good on your golden set. Keep the old index for a soak period.
  • Estimating the migration: multiply documents by tokens per document. Fifty million documents of 500 tokens is 25 billion tokens: about 7 hours at a million tokens per second across your workers, and \$500 at \$0.02 per million tokens (replace all three inputs with your own). The money is usually modest; the time, the rate limits and the double storage are what need planning.
  • The model doesn't know your jargon: measure on your own questions first, because public leaderboards measure public data. Add keyword search for exact terms, then fine-tune on domain pairs if the gap remains.
  • A wrong answer: ask one question first. Was a relevant document in the retrieved top k? If not, it is retrieval: chunking, hybrid search, reranking, the model. If yes and the answer ignored it, it is generation: the prompt, the context order, fewer distracting chunks.
  • Whatever you do, keep the golden set and rerun it on every change. It is the only number that predicts your system's quality.

What it costs. A golden set: a few hundred questions with their answer documents, taken from resolved support or search sessions, which cost nothing to collect. Re-embedding: linear in the corpus, so ten times the documents means ten times the hours and the dollars, plus double storage while both indexes exist. Blue/green adds one extra write per new document during the migration and one shadow evaluation. Fine-tuning adds labeled pairs and a training run, and then a migration, because a fine-tuned model is a new map. Triage costs a person reading a handful of failures; only the retrieval half can be measured without a model, which is why recall@k on the golden set is the first number to track.

What breaks.

  • Mixed spaces. Documents from one model, queries from another, the same number of dimensions: recall falls from 0.92 to 0.25 with no error. A model with a different number of dimensions at least crashes. Tie every index to a model version and refuse vectors from any other.
  • A backfill that never finished. Only new documents got re-embedded. Dual-writing and backfilling are separate steps, and the migration is not done until every old document has been re-embedded.
  • Cutting over on faith. The shadow comparison is the gate: this lesson's migration refuses a candidate that scores 0.58 against the live index's 0.92.
  • No way back. Retire the old index only after a soak period; until then, rollback is repointing the alias.
  • Trusting a leaderboard. BEIR showed that retrieval models trained on one domain often lose badly on others, and MTEB found no single model wins everywhere. Your questions are the benchmark.
  • Tuning the prompt for a retrieval failure. The error-code question in this lesson stays wrong at every k because its page ranks fourth; no prompt fixes that. Retrieving more turns some retrieval failures into generation failures, so raising k is not a fix either.

In the wild. Index aliases are the switch: Elasticsearch's aliases API swaps an alias from one index to another in a single atomic operation with no downtime, and Qdrant's collection aliases let you build a second collection in the background and switch atomically with no concurrent request affected. Qdrant also fixes the vector size per collection, so a model with a new number of dimensions is a new collection by construction. Martin Fowler's description of blue/green deployment is where the pattern's name comes from. The MTEB leaderboard on Hugging Face ranks embedding models across tasks, and BEIR is the zero-shot retrieval benchmark that showed how much domain matters. Ragas separates retrieval metrics (context precision, context recall) from generation metrics (faithfulness, response relevancy), the same split as this lesson's triage. The papers behind this lesson are listed at the end.

Go deeper. Level 2 measures the mixed-space failure on the golden set, walks a blue/green migration state by state with the code that refuses a worse model, works the re-embedding arithmetic, shows the jargon gap and a tuned model closing it, and runs every golden question through a triage that names the failing half. If you only needed the runbook, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

Level 2 makes each of those truths concrete, starting with two mapmakers.

The everyday picture. Two mapmakers each draw a map of the same city with their own grid. On one map, square (3, 7) is the train station; on the other, (3, 7) is a park. Both maps are fine, but you can't read a position off one map and look it up on the other.

Every embedding model is its own mapmaker. Upgrade the model and every document has to be re-plotted on the new map, and until that's done, the old map and the new one must never be mixed. That's the first of three operational truths this lesson makes concrete:

  1. Changing models means re-embedding everything, and doing it without downtime or a silent quality drop takes a plan.
  2. General models don't speak your company's jargon, so measure on your own questions and adapt when needed.
  3. When a retrieval-augmented system answers wrongly, find out which half failed: retrieval (the right page was never found) or generation (it was found and misused).

Every model has its own space

Tiny worked example: the word "vpn" embedded by two versions of the toy model (primer.common.embedder) with the same 64 dimensions: their cosine is about 0.06, essentially unrelated, though it's the same word. Now take 12 real questions with known answers (the "golden set" in primer.common.corpus) over 20 documents. Search v1 documents with v1 queries and the right answer is in the top 3 for 11 of 12 questions (0.92). Search the same v1 documents with queries from the new v2 model and it drops to 0.25, about what random ranking would give (3 of 20 documents shown, so 15%), and nothing crashes.

Level 3: the formula and its symbols

$$ \text{hit@}k = \frac{1}{|Q|} \sum_{q \in Q} \big[\, \text{rel}(q) \cap \text{top}_k(q) \neq \varnothing \,\big] $$

Symbols

Symbol Meaning here Example
Q the golden set of questions 12 questions
|Q| how many questions 12
q ∈ Q each question in turn
rel(q) the documents that truly answer q {it-004}
top_k(q) the k documents search returned 3 documents
∩, ≠ ∅ "they share at least one document"
[ … ] 1 if true, 0 if false
hit@k the share of questions with a right answer in the top k 0 to 1

In words: the share of golden questions for which at least one right document appears among the top k results. (With one relevant document per question, as here, this equals recall@k.)

On the example: 11 hits out of 12 = 0.92 with matching models; 3 out of 12 = 0.25 with mixed models.

Level 3: in Python

In Python:

def hit(rel_q, top_q):
    # [rel(q) ∩ top_k(q) ≠ ∅]
    return 1 if rel_q & top_q else 0
hit({"it-004"}, {"it-002", "it-004", "fin-001"})  # → 1
# matching models: 11 of the 12 questions
hits = [1] * 11 + [0]
# (1/|Q|) Σ over q in Q
round(sum(hits) / len(hits), 2)  # → 0.92
# mixed models: 3 of the 12
hits = [1] * 3 + [0] * 9
round(sum(hits) / len(hits), 2)  # → 0.25

Recall@3 is 0.92 when documents and queries share a model, and falls to 0.25, near the 0.15 of random ranking, when v2 queries search v1 documents

Reading it: each bar is recall@3 on the golden set. The first two bars use one model for both documents and queries, v1 then v2: both work. The third bar searches v1 documents with v2 queries: recall falls to 0.25 (3 of the 12 questions), barely above the dashed line at 0.15, which is what random ranking would score (3 of 20 documents shown). A model with a different number of dimensions would at least crash (you can't dot a 128-number vector with a 64-number one); a same-size model fails silently, which is worse.

In code: golden_recall embeds the documents with one model and the questions with another and returns hit@k. VectorIndex is a flat index tied to one model version: VectorIndex.search embeds a text query, and VectorIndex.search_vector takes a vector that is already made.

Why it matters: "we upgraded the embedding model and search got weird" is a classic incident. The cause is almost always documents and queries embedded by different models, typically because only new documents were re-embedded.

Migrating without a bad day: blue/green

Everyday picture: building a new bridge next to the old one. Traffic keeps using the old bridge while the new one is built and inspected. When the new bridge passes inspection, traffic is switched over, and the old bridge stays standing for a while in case something turns up.

stateDiagram-v2 [*] --> Live_v1 Live_v1 --> DualWrite: start v2 (empty index, new docs go to both) DualWrite --> Backfilled: backfill (re-embed old docs in batches) Backfilled --> Compared: shadow compare (recall on golden set) Compared --> Live_v2: cutover, only if v2 at least as good Compared --> Live_v1: refused, v2 is worse Live_v2 --> Live_v1: rollback (v1 index kept) Live_v2 --> [*]: retire v1 after a soak period

Reading it: follow the states from the top. Users are served by v1 the whole time until cutover. Dual-writing from the very first step means documents added mid-migration land in both indexes, so the new one never falls behind. The shadow comparison is the gate: the switch only happens if v2 is at least as good on your own golden set (BlueGreenMigration refuses to cut over without it). The old index is kept after cutover, so rollback is flipping one pointer, not a multi-hour rebuild.

The "live alias" is a pointer: the search service asks for "the live index", and cutover or rollback just repoints it. Many vector databases support aliases for exactly this.

In code: each arrow of the diagram is one method: BlueGreenMigration.start, BlueGreenMigration.add (the dual write), BlueGreenMigration.backfill, BlueGreenMigration.shadow_compare, BlueGreenMigration.cutover and BlueGreenMigration.rollback.

Tiny worked example: what will re-embedding cost? 50 million documents of about 500 tokens each, an embedding throughput of 1 million tokens per second across your workers, and a price of \$0.02 per million tokens (all three are inputs you replace with your own numbers):

Level 3: the formula and its symbols

$$ \text{hours} = \frac{n \cdot t}{r \cdot 3600}, \qquad \text{dollars} = \frac{n \cdot t}{10^6} \cdot p $$

Symbols

Symbol Meaning here Example
n number of documents (or chunks) 50,000,000
t average tokens per document 500
n · t total tokens to embed 25,000,000,000
r throughput, tokens per second 1,000,000
3600 seconds per hour
p price per million tokens \$0.02

In words: total tokens divided by throughput gives the time; total tokens in millions times the price gives the cost.

On the example: 25 × 10⁹ / (10⁶ × 3600) = 6.94 hours; 25,000 × \$0.02 = \$500.

Level 3: in Python

In Python:

# documents, tokens per document
n, t = 50_000_000, 500
# tokens per second, dollars per million tokens
r, p = 1_000_000, 0.02
# total tokens
n * t  # → 25000000000
# hours
round(n * t / (r * 3600), 2)  # → 6.94
# dollars
round(n * t / 10**6 * p, 2)  # → 500.0

Re-embedding time and cost both grow in step with corpus size: ten times the documents, ten times the hours and the dollars

Reading it: the horizontal axis is corpus size (log scale); the left panel shows hours at three throughputs, the right panel shows dollars at three prices. Both grow in straight lines on these log axes: ten times the documents, ten times the time and money. The money is usually modest; the time, the rate limits and the double storage during the migration are what need planning.

In code: reembed_estimate computes the tokens, hours and dollars for your own n, t, r and p.

Why it matters: re-embedding is routine: new models, fine-tunes, chunking changes and bug fixes all require it. Versioned indexes, dual writes, a golden-set gate and a rollback path turn it from a risky event into a boring one.

Domain mismatch: fluent, but not in your jargon

Everyday picture: a new hire who speaks perfect English but doesn't yet know that "AP" means accounts payable or that "T&E" means travel and expenses. General embedding models are trained mostly on web text; your acronyms and product names are new words to them.

Tiny worked example: four real-sounding employee questions in company jargon: "ap aging report", "hotspot from a client site", "sso lockout", "t-and-e submission deadline". The general toy model puts the right document first for 1 of 4. A version that has learned the four jargon words (DomainTunedEmbedder, standing in for a model fine-tuned on company pairs) gets 4 of 4.

On company jargon the general model buries three of the four right answers, while the tuned model ranks every one first

Reading it: each pair of bars is one jargon question; the height is where the right document ranked (1 is best, shorter is better). The general model buries three of the four answers; the tuned model puts every one first. The one the general model gets right, "sso lockout", is saved by a word it does know ("lockout").

In code: rank_of_first_relevant finds where the right document lands for one question under one model: the height of each bar.

flowchart LR L["Search & support logs"] --> F["Keep resolved sessions<br/>(the clicked doc answered it)"] F --> N["Normalize and merge<br/>repeated questions"] N --> G["Golden set: question → relevant docs<br/>(a few hundred is plenty)"] G --> T["Test several models on YOUR set"] T --> D{Good enough?} D -->|no| FT["Fine-tune on domain pairs,<br/>add hybrid search"] D -->|yes| S[Ship, and keep the set for regressions] FT --> T

Reading it: the evaluation set comes from real usage, not invention. Resolved sessions give you (question, answer document) pairs for free; build_eval_set keeps only resolved ones and merges repeats. Test candidate models on that set, and when none is good enough, fine-tune on domain pairs (primer.ml.embeddings.contrastive) and add keyword search, which matches jargon exactly (primer.ml.embeddings.retrieval).

Why it matters: public leaderboards measure public data. The only number that predicts your system's quality is recall on your own questions.

Measure retrieval separately from generation

Everyday picture: an open-book exam. A wrong answer has one of two causes: the right page wasn't in the book you brought (retrieval failure), or it was, and you misread it (generation failure). Studying harder fixes the second, not the first.

Tiny worked example: three answered questions, checked by hand.

Retrieved (top 3) Truly relevant Answer used Verdict
it-002, fin-006, fin-001 it-004 it-002 retrieval failure: it-004 was never found
hr-002, hr-001, hr-005 hr-001 hr-002 generation failure: found but not used
it-001, it-002 it-001 it-001 correct
flowchart TD A[Wrong answer] --> R{Was a relevant document<br/>in the retrieved top k?} R -->|no| RF[Retrieval failure:<br/>fix chunking, hybrid search,<br/>reranking, the embedding model] R -->|yes| G{Did the answer use it?} G -->|no| GF[Generation failure:<br/>fix the prompt, context order,<br/>fewer distracting chunks] G -->|yes| OK[Check the grader:<br/>the answer may be fine]

Reading it: one question splits every failure in two. If the right document wasn't retrieved, no prompt change can help, so work on retrieval. If it was retrieved and ignored, the problem is downstream, in the prompt or the context. Only the retrieval half can be measured without a model, which is why recall@k on a golden set is the first number to track.

Retrieving more documents turns some retrieval failures into generation failures, and the error-code question stays wrong at every k

Reading it: each bar is the 12 golden questions, answered by a toy generator that always uses the top document, with k documents retrieved. Green is correct, red is a retrieval failure, orange is a generation failure. Retrieving more (larger k) turns some retrieval failures into generation failures: the right page is now in the book, but the reader still opened the wrong one. Watch the error-code question ("what does ERR-4012 mean"): its answer ranks only fourth, so it's a retrieval failure until k = 5, and even then the top document is the wrong one. Dense vectors blur exact identifiers; keyword search would put that page first.

In code: triage gives the verdict for one answered question, following the flowchart above; triage_golden_set runs every golden question through retrieval and a toy generator that answers from the top document.

Why it matters: teams burn weeks tuning prompts for failures that were retrieval all along. Triage first, then fix the half that's broken.

In 20 seconds

  • Every embedding model has its own vector space; never mix documents and queries from different models. Same-size models fail silently.
  • Migrate blue/green: new index, dual-write, backfill, compare on a golden set, cut over only if better, keep the old index for rollback.
  • Re-embedding cost is n × tokens: estimate hours and dollars before you start.
  • General models miss company jargon; build an eval set from resolved logs, and fine-tune or add keyword search when needed.
  • Triage RAG failures into retrieval vs. generation before fixing anything.

Self-test questions

Q: You need to switch embedding models on a 50-million-document index. Plan the migration. Create a versioned index for the new model and dual-write new documents to both. Backfill by re-embedding existing documents in restartable batches (estimate: 50M × 500 tokens = 25B tokens; at 1M tokens/s that's about 7 hours). Compare recall on a golden set in shadow; cut over the live alias only if the new model is at least as good; keep the old index for fast rollback, then retire it.

Q: Why can't you compare vectors from two different embedding models? Each model defines its own coordinate system. Even with the same number of dimensions, the same text lands in unrelated places, so similarity between the two spaces is meaningless.

Q: A general embedding model performs poorly on a company's internal documents. What helps? Build a small eval set from real queries and the documents that resolved them, test several models on it, add hybrid (keyword + dense) search for exact terms, and fine-tune an embedding model on domain pairs with hard negatives if needed.

Q: Your RAG system gives confident wrong answers. What's your first diagnostic step? Split retrieval from generation: for failing questions, check whether a relevant document was in the retrieved top k. If not, it's retrieval (fix search); if yes, it's generation (fix prompt and context). Track recall@k on a labeled set continuously.

The papers behind this lesson

  • Thakur et al., BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models (2021): https://arxiv.org/abs/2104.08663. Showed that retrieval models trained on one domain often lose badly on others, which is why you evaluate on your own data.
  • Muennighoff et al., MTEB: Massive Text Embedding Benchmark (2022): https://arxiv.org/abs/2210.07316. Compared embedding models across many tasks and found no single model wins everywhere.

Further reading

on GitHub
  1r"""
  2# Operating embeddings: model changes, jargon, and finding what broke
  3
  4Run: `python -m primer.ml.embeddings.operations`
  5
  6New to vectors or Σ? `primer.notation` builds them from zero. This lesson
  7builds on `primer.ml.embeddings.similarity` and `primer.ml.embeddings.contrastive`.
  8
  9## Level 1: The practitioner's guide
 10
 11**In one sentence.** Operating embeddings means keeping every query on the
 12same map as the documents it searches, changing that map without a bad day,
 13checking that the map speaks your domain's language, and knowing which half
 14of a retrieval-augmented system failed when an answer is wrong.
 15
 16**When you need it.** From the day an embedding index serves real traffic.
 17Every embedding model is its own coordinate system: in this lesson, the word
 18"vpn" embedded by two versions of the same toy model, both 64 dimensions,
 19has a cosine similarity of about 0.05, as if they were unrelated words.
 20Search documents embedded by one model with queries embedded by another and
 21recall@3 on this lesson's golden set falls from 0.92 to 0.25, barely above
 22the 0.15 that random ranking would score, and nothing crashes. You will
 23change models: new releases, a fine-tune, a chunking change and a bug fix
 24all require re-embedding everything. You will meet jargon the model doesn't
 25know: this lesson's general model puts the right document first for one of
 26four company-jargon questions. And you will get a confident wrong answer and
 27need to know whether retrieval or generation caused it. The tell: "we
 28upgraded the embedding model and search got weird", or weeks spent tuning
 29prompts for answers whose right page was never retrieved.
 30
 31**Your options.** Two decisions recur: how to change the model, and what to
 32do when it doesn't know your words.
 33
 34| Option | What it does | What it gives you | What it costs | Where it lives |
 35|---|---|---|---|---|
 36| Re-embed only new documents | Old documents keep their old vectors; new ones get the new model | Nothing: it is the classic incident. Queries and documents stop sharing a space and recall collapses silently | The cheapest to run and the most expensive to discover | A pipeline that forgot the backfill |
 37| Re-embed in place | Stops writes, re-embeds every document into the same index, resumes | A correct index at the end | Downtime or a window of mixed results, and no rollback except re-embedding again | Your ingestion job |
 38| Blue/green migration | Builds a second index alongside, dual-writes new documents to both, backfills, compares on a golden set, then repoints a live alias | No downtime, a gate that refuses a worse model, rollback by flipping one pointer | Double storage during the migration, a golden set, the backfill time | Your code plus index aliases (Elasticsearch, Qdrant) |
 39| Keep the general model, add keyword search | Fuses BM25 with dense search so exact jargon matches by spelling | Identifiers and acronyms found without training | Two searches per query | Any search engine (`primer.ml.embeddings.retrieval`) |
 40| Fine-tune on domain pairs | Trains the embedding model on your own (question, document) pairs | Four of four jargon questions right in this lesson, against one of four | Labeled pairs, a training run, then a full re-embed | `primer.ml.embeddings.contrastive`, followed by a migration |
 41| Measure and triage | Keeps a golden set of questions with known answer passages and checks it on every change | Silent failures made visible, each split into retrieval or generation | A few hundred labeled questions, taken from resolved logs | Your evaluation harness; Ragas |
 42
 43**How to choose.**
 44
 45- Changing the model on a live index: blue/green, every time. Version the
 46  index by model, dual-write from the first minute so nothing falls behind,
 47  backfill in restartable batches, and cut over only when the new index is
 48  at least as good on your golden set. Keep the old index for a soak period.
 49- Estimating the migration: multiply documents by tokens per document.
 50  Fifty million documents of 500 tokens is 25 billion tokens: about 7 hours
 51  at a million tokens per second across your workers, and \$500 at \$0.02
 52  per million tokens (replace all three inputs with your own). The money is
 53  usually modest; the time, the rate limits and the double storage are what
 54  need planning.
 55- The model doesn't know your jargon: measure on your own questions first,
 56  because public leaderboards measure public data. Add keyword search for
 57  exact terms, then fine-tune on domain pairs if the gap remains.
 58- A wrong answer: ask one question first. Was a relevant document in the
 59  retrieved top k? If not, it is retrieval: chunking, hybrid search,
 60  reranking, the model. If yes and the answer ignored it, it is generation:
 61  the prompt, the context order, fewer distracting chunks.
 62- Whatever you do, keep the golden set and rerun it on every change. It is
 63  the only number that predicts your system's quality.
 64
 65**What it costs.** A golden set: a few hundred questions with their answer
 66documents, taken from resolved support or search sessions, which cost
 67nothing to collect. Re-embedding: linear in the corpus, so ten times the
 68documents means ten times the hours and the dollars, plus double storage
 69while both indexes exist. Blue/green adds one extra write per new document
 70during the migration and one shadow evaluation. Fine-tuning adds labeled
 71pairs and a training run, and then a migration, because a fine-tuned model
 72is a new map. Triage costs a person reading a handful of failures; only the
 73retrieval half can be measured without a model, which is why recall@k on
 74the golden set is the first number to track.
 75
 76**What breaks.**
 77
 78- **Mixed spaces.** Documents from one model, queries from another, the same
 79  number of dimensions: recall falls from 0.92 to 0.25 with no error. A
 80  model with a different number of dimensions at least crashes. Tie every
 81  index to a model version and refuse vectors from any other.
 82- **A backfill that never finished.** Only new documents got re-embedded.
 83  Dual-writing and backfilling are separate steps, and the migration is not
 84  done until every old document has been re-embedded.
 85- **Cutting over on faith.** The shadow comparison is the gate: this
 86  lesson's migration refuses a candidate that scores 0.58 against the live
 87  index's 0.92.
 88- **No way back.** Retire the old index only after a soak period; until
 89  then, rollback is repointing the alias.
 90- **Trusting a leaderboard.** BEIR showed that retrieval models trained on
 91  one domain often lose badly on others, and MTEB found no single model
 92  wins everywhere. Your questions are the benchmark.
 93- **Tuning the prompt for a retrieval failure.** The error-code question in
 94  this lesson stays wrong at every k because its page ranks fourth; no
 95  prompt fixes that. Retrieving more turns some retrieval failures into
 96  generation failures, so raising k is not a fix either.
 97
 98**In the wild.** Index aliases are the switch: Elasticsearch's aliases API
 99swaps an alias from one index to another in a single atomic operation with
100no downtime, and Qdrant's collection aliases let you build a second
101collection in the background and switch atomically with no concurrent
102request affected. Qdrant also fixes the vector size per collection, so a
103model with a new number of dimensions is a new collection by construction.
104Martin Fowler's description of blue/green deployment is where the pattern's
105name comes from. The MTEB leaderboard on Hugging Face ranks embedding models
106across tasks, and BEIR is the zero-shot retrieval benchmark that showed how
107much domain matters. Ragas separates retrieval metrics (context precision,
108context recall) from generation metrics (faithfulness, response relevancy),
109the same split as this lesson's triage. The papers behind this lesson are
110listed at the end.
111
112**Go deeper.** Level 2 measures the mixed-space failure on the golden set,
113walks a blue/green migration state by state with the code that refuses a
114worse model, works the re-embedding arithmetic, shows the jargon gap and a
115tuned model closing it, and runs every golden question through a triage
116that names the failing half. If you only needed the runbook, you are done.
117
118## Level 2: How it works, from scratch
119
120Level 2 makes each of those truths concrete, starting with two mapmakers.
121
122**The everyday picture.** Two mapmakers each draw a map of the same city with their own grid. On one
123map, square (3, 7) is the train station; on the other, (3, 7) is a park. Both
124maps are fine, but you can't read a position off one map and look it up on
125the other.
126
127Every embedding model is its own mapmaker. Upgrade the model and every
128document has to be re-plotted on the new map, and until that's done, the old
129map and the new one must never be mixed. That's the first of three
130operational truths this lesson makes concrete:
131
1321. **Changing models means re-embedding everything**, and doing it without
133   downtime or a silent quality drop takes a plan.
1342. **General models don't speak your company's jargon**, so measure on your
135   own questions and adapt when needed.
1363. **When a retrieval-augmented system answers wrongly, find out which half
137   failed**: retrieval (the right page was never found) or generation (it was
138   found and misused).
139
140## Every model has its own space
141
142**Tiny worked example:** the word "vpn" embedded by two versions of the toy
143model (`primer.common.embedder`) with the same 64 dimensions: their cosine
144is about 0.06, essentially unrelated, though it's the same word. Now take 12
145real questions with known answers (the "golden set" in `primer.common.corpus`)
146over 20 documents. Search v1 documents with v1 queries and the right answer
147is in the top 3 for 11 of 12 questions (**0.92**). Search the *same* v1
148documents with queries from the new v2 model and it drops to **0.25**, about
149what random ranking would give (3 of 20 documents shown, so 15%), and
150**nothing crashes**.
151
152$$
153\text{hit@}k = \frac{1}{|Q|} \sum_{q \in Q} \big[\, \text{rel}(q) \cap \text{top}_k(q) \neq \varnothing \,\big]
154$$
155
156**Symbols**
157
158| Symbol | Meaning here | Example |
159|---|---|---|
160| Q | the golden set of questions | 12 questions |
161| \|Q\| | how many questions | 12 |
162| q ∈ Q | each question in turn | |
163| rel(q) | the documents that truly answer q | {it-004} |
164| top_k(q) | the k documents search returned | 3 documents |
165| ∩, ≠ ∅ | "they share at least one document" | |
166| [ … ] | 1 if true, 0 if false | |
167| hit@k | the share of questions with a right answer in the top k | 0 to 1 |
168
169**In words:** the share of golden questions for which at least one right
170document appears among the top k results. (With one relevant document per
171question, as here, this equals recall@k.)
172
173**On the example:** 11 hits out of 12 = 0.92 with matching models; 3 out of 12 =
1740.25 with mixed models.
175
176**In Python:**
177
178```python
179def hit(rel_q, top_q):
180    # [rel(q) ∩ top_k(q) ≠ ∅]
181    return 1 if rel_q & top_q else 0
182hit({"it-004"}, {"it-002", "it-004", "fin-001"})  # → 1
183# matching models: 11 of the 12 questions
184hits = [1] * 11 + [0]
185# (1/|Q|) Σ over q in Q
186round(sum(hits) / len(hits), 2)  # → 0.92
187# mixed models: 3 of the 12
188hits = [1] * 3 + [0] * 9
189round(sum(hits) / len(hits), 2)  # → 0.25
190```
191
192![Recall@3 is 0.92 when documents and queries share a model, and falls to 0.25, near the 0.15 of random ranking, when v2 queries search v1 documents](figures/primer.ml.embeddings.operations.cross_model.svg)
193
194**Reading it:** each bar is recall@3 on the golden set. The first two bars
195use one model for both documents and queries, v1 then v2: both work. The
196third bar searches v1 documents with v2 queries: recall falls to 0.25 (3 of
197the 12 questions), barely above the dashed line at 0.15, which is what random
198ranking would score (3 of 20 documents shown). A model with a *different*
199number of dimensions would at least crash (you can't dot a 128-number vector
200with a 64-number one); a same-size model fails silently, which is worse.
201
202**In code:** `golden_recall` embeds the documents with one model and the
203questions with another and returns hit@k. `VectorIndex` is a flat index tied
204to one model version: `VectorIndex.search` embeds a text query, and
205`VectorIndex.search_vector` takes a vector that is already made.
206
207**Why it matters:** "we upgraded the embedding model and search got weird"
208is a classic incident. The cause is almost always documents and queries
209embedded by different models, typically because only new documents were
210re-embedded.
211
212## Migrating without a bad day: blue/green
213
214**Everyday picture:** building a new bridge next to the old one. Traffic
215keeps using the old bridge while the new one is built and inspected. When
216the new bridge passes inspection, traffic is switched over, and the old
217bridge stays standing for a while in case something turns up.
218
219```mermaid
220stateDiagram-v2
221  [*] --> Live_v1
222  Live_v1 --> DualWrite: start v2 (empty index, new docs go to both)
223  DualWrite --> Backfilled: backfill (re-embed old docs in batches)
224  Backfilled --> Compared: shadow compare (recall on golden set)
225  Compared --> Live_v2: cutover, only if v2 at least as good
226  Compared --> Live_v1: refused, v2 is worse
227  Live_v2 --> Live_v1: rollback (v1 index kept)
228  Live_v2 --> [*]: retire v1 after a soak period
229```
230
231**Reading it:** follow the states from the top. Users are served by v1 the
232whole time until cutover. Dual-writing from the very first step means
233documents added mid-migration land in both indexes, so the new one never
234falls behind. The **shadow comparison** is the gate: the switch only happens
235if v2 is at least as good on your own golden set (`BlueGreenMigration`
236refuses to cut over without it). The old index is kept after cutover, so
237rollback is flipping one pointer, not a multi-hour rebuild.
238
239The "live alias" is a pointer: the search service asks for "the live
240index", and cutover or rollback just repoints it. Many vector databases
241support aliases for exactly this.
242
243**In code:** each arrow of the diagram is one method: `BlueGreenMigration.start`,
244`BlueGreenMigration.add` (the dual write), `BlueGreenMigration.backfill`,
245`BlueGreenMigration.shadow_compare`, `BlueGreenMigration.cutover` and
246`BlueGreenMigration.rollback`.
247
248**Tiny worked example: what will re-embedding cost?** 50 million documents of
249about 500 tokens each, an embedding throughput of 1 million tokens per second
250across your workers, and a price of \$0.02 per million tokens (all three are
251inputs you replace with your own numbers):
252
253$$
254\text{hours} = \frac{n \cdot t}{r \cdot 3600}, \qquad
255\text{dollars} = \frac{n \cdot t}{10^6} \cdot p
256$$
257
258**Symbols**
259
260| Symbol | Meaning here | Example |
261|---|---|---|
262| n | number of documents (or chunks) | 50,000,000 |
263| t | average tokens per document | 500 |
264| n · t | total tokens to embed | 25,000,000,000 |
265| r | throughput, tokens per second | 1,000,000 |
266| 3600 | seconds per hour | |
267| p | price per million tokens | \$0.02 |
268
269**In words:** total tokens divided by throughput gives the time; total
270tokens in millions times the price gives the cost.
271
272**On the example:** 25 × 10⁹ / (10⁶ × 3600) = **6.94 hours**; 25,000 × \$0.02 =
273**\$500**.
274
275**In Python:**
276
277```python
278# documents, tokens per document
279n, t = 50_000_000, 500
280# tokens per second, dollars per million tokens
281r, p = 1_000_000, 0.02
282# total tokens
283n * t  # → 25000000000
284# hours
285round(n * t / (r * 3600), 2)  # → 6.94
286# dollars
287round(n * t / 10**6 * p, 2)  # → 500.0
288```
289
290![Re-embedding time and cost both grow in step with corpus size: ten times the documents, ten times the hours and the dollars](figures/primer.ml.embeddings.operations.reembed.svg)
291
292**Reading it:** the horizontal axis is corpus size (log scale); the left
293panel shows hours at three throughputs, the right panel shows dollars at
294three prices. Both grow in straight lines on these log axes: ten times the
295documents, ten times the time and money. The money is usually modest; the
296time, the rate limits and the double storage during the migration are what
297need planning.
298
299**In code:** `reembed_estimate` computes the tokens, hours and dollars for
300your own n, t, r and p.
301
302**Why it matters:** re-embedding is routine: new models, fine-tunes,
303chunking changes and bug fixes all require it. Versioned indexes, dual
304writes, a golden-set gate and a rollback path turn it from a risky event
305into a boring one.
306
307## Domain mismatch: fluent, but not in your jargon
308
309**Everyday picture:** a new hire who speaks perfect English but doesn't yet
310know that "AP" means accounts payable or that "T&E" means travel and
311expenses. General embedding models are trained mostly on web text; your
312acronyms and product names are new words to them.
313
314**Tiny worked example:** four real-sounding employee questions in company
315jargon: "ap aging report", "hotspot from a client site", "sso lockout",
316"t-and-e submission deadline". The general toy model puts the right
317document first for **1 of 4**. A version that has learned the four jargon
318words (`DomainTunedEmbedder`, standing in for a model fine-tuned on
319company pairs) gets **4 of 4**.
320
321![On company jargon the general model buries three of the four right answers, while the tuned model ranks every one first](figures/primer.ml.embeddings.operations.jargon.svg)
322
323**Reading it:** each pair of bars is one jargon question; the height is where
324the right document ranked (1 is best, shorter is better). The general model
325buries three of the four answers; the tuned model puts every one first. The
326one the general model gets right, "sso lockout", is saved by a word it does
327know ("lockout").
328
329**In code:** `rank_of_first_relevant` finds where the right document lands
330for one question under one model: the height of each bar.
331
332```mermaid
333flowchart LR
334  L["Search & support logs"] --> F["Keep resolved sessions<br/>(the clicked doc answered it)"]
335  F --> N["Normalize and merge<br/>repeated questions"]
336  N --> G["Golden set: question → relevant docs<br/>(a few hundred is plenty)"]
337  G --> T["Test several models on YOUR set"]
338  T --> D{Good enough?}
339  D -->|no| FT["Fine-tune on domain pairs,<br/>add hybrid search"]
340  D -->|yes| S[Ship, and keep the set for regressions]
341  FT --> T
342```
343
344**Reading it:** the evaluation set comes from real usage, not invention.
345Resolved sessions give you (question, answer document) pairs for free;
346`build_eval_set` keeps only resolved ones and merges repeats. Test
347candidate models on that set, and when none is good enough, fine-tune on
348domain pairs (`primer.ml.embeddings.contrastive`) and add keyword search,
349which matches jargon exactly (`primer.ml.embeddings.retrieval`).
350
351**Why it matters:** public leaderboards measure public data. The only
352number that predicts your system's quality is recall on your own
353questions.
354
355## Measure retrieval separately from generation
356
357**Everyday picture:** an open-book exam. A wrong answer has one of two
358causes: the right page wasn't in the book you brought (**retrieval
359failure**), or it was, and you misread it (**generation failure**). Studying
360harder fixes the second, not the first.
361
362**Tiny worked example:** three answered questions, checked by hand.
363
364| Retrieved (top 3) | Truly relevant | Answer used | Verdict |
365|---|---|---|---|
366| it-002, fin-006, fin-001 | it-004 | it-002 | retrieval failure: it-004 was never found |
367| hr-002, hr-001, hr-005 | hr-001 | hr-002 | generation failure: found but not used |
368| it-001, it-002 | it-001 | it-001 | correct |
369
370```mermaid
371flowchart TD
372  A[Wrong answer] --> R{Was a relevant document<br/>in the retrieved top k?}
373  R -->|no| RF[Retrieval failure:<br/>fix chunking, hybrid search,<br/>reranking, the embedding model]
374  R -->|yes| G{Did the answer use it?}
375  G -->|no| GF[Generation failure:<br/>fix the prompt, context order,<br/>fewer distracting chunks]
376  G -->|yes| OK[Check the grader:<br/>the answer may be fine]
377```
378
379**Reading it:** one question splits every failure in two. If the right
380document wasn't retrieved, no prompt change can help, so work on retrieval.
381If it was retrieved and ignored, the problem is downstream, in the prompt or
382the context. Only the retrieval half can be measured without a model, which
383is why recall@k on a golden set is the first number to track.
384
385![Retrieving more documents turns some retrieval failures into generation failures, and the error-code question stays wrong at every k](figures/primer.ml.embeddings.operations.triage.svg)
386
387**Reading it:** each bar is the 12 golden questions, answered by a toy
388generator that always uses the top document, with k documents retrieved.
389Green is correct, red is a retrieval failure, orange is a generation
390failure. Retrieving more (larger k) turns some retrieval failures into
391generation failures: the right page is now in the book, but the reader still
392opened the wrong one. Watch the error-code question ("what does ERR-4012
393mean"): its answer ranks only fourth, so it's a retrieval failure until k = 5,
394and even then the top document is the wrong one. Dense vectors blur exact
395identifiers; keyword search would put that page first.
396
397**In code:** `triage` gives the verdict for one answered question, following
398the flowchart above; `triage_golden_set` runs every golden question through
399retrieval and a toy generator that answers from the top document.
400
401**Why it matters:** teams burn weeks tuning prompts for failures that were
402retrieval all along. Triage first, then fix the half that's broken.
403
404## In 20 seconds
405- Every embedding model has its own vector space; never mix documents and
406  queries from different models. Same-size models fail silently.
407- Migrate blue/green: new index, dual-write, backfill, compare on a golden
408  set, cut over only if better, keep the old index for rollback.
409- Re-embedding cost is n × tokens: estimate hours and dollars before you start.
410- General models miss company jargon; build an eval set from resolved logs,
411  and fine-tune or add keyword search when needed.
412- Triage RAG failures into retrieval vs. generation before fixing anything.
413
414## Self-test questions
415
416**Q: You need to switch embedding models on a 50-million-document index. Plan the migration.**
417Create a versioned index for the new model and dual-write new documents to
418both. Backfill by re-embedding existing documents in restartable batches
419(estimate: 50M × 500 tokens = 25B tokens; at 1M tokens/s that's about 7 hours).
420Compare recall on a golden set in shadow; cut over the live alias only if the
421new model is at least as good; keep the old index for fast rollback, then
422retire it.
423
424**Q: Why can't you compare vectors from two different embedding models?**
425Each model defines its own coordinate system. Even with the same number of
426dimensions, the same text lands in unrelated places, so similarity between
427the two spaces is meaningless.
428
429**Q: A general embedding model performs poorly on a company's internal documents. What helps?**
430Build a small eval set from real queries and the documents that resolved
431them, test several models on it, add hybrid (keyword + dense) search for
432exact terms, and fine-tune an embedding model on domain pairs with hard
433negatives if needed.
434
435**Q: Your RAG system gives confident wrong answers. What's your first diagnostic step?**
436Split retrieval from generation: for failing questions, check whether a
437relevant document was in the retrieved top k. If not, it's retrieval (fix
438search); if yes, it's generation (fix prompt and context). Track recall@k on
439a labeled set continuously.
440
441## The papers behind this lesson
442
443- **Thakur et al., *BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models* (2021)**: https://arxiv.org/abs/2104.08663.
444  Showed that retrieval models trained on one domain often lose badly on others, which is why you evaluate on your own data.
445- **Muennighoff et al., *MTEB: Massive Text Embedding Benchmark* (2022)**: https://arxiv.org/abs/2210.07316.
446  Compared embedding models across many tasks and found no single model wins everywhere.
447
448## Further reading
449- Es et al., *RAGAS: Automated Evaluation of Retrieval Augmented Generation* (2023): https://arxiv.org/abs/2309.15217
450- Martin Fowler, *BlueGreenDeployment*: https://martinfowler.com/bliki/BlueGreenDeployment.html
451- MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard
452"""
453
454from __future__ import annotations
455
456from dataclasses import dataclass, field
457
458import numpy as np
459
460from primer._show import banner, say, table, takeaway
461from primer.common.corpus import DOCS, LABELED_QUERIES, Doc
462from primer.common.embedder import ConceptEmbedder, _unit_vector
463from primer.common.text import tokenize
464
465# Model versions. Each `name` is a different model, so each defines its own
466# vector space, even when two have the same number of dimensions.
467CURRENT_MODEL = ConceptEmbedder()  # the repo's standard toy model, "concept-v1", 64 dims
468SAME_SIZE_NEW_MODEL = ConceptEmbedder(name="concept-v2")  # also 64 dims: mixing it with v1 fails *silently*
469BETTER_MODEL = ConceptEmbedder(dim=128, name="concept-v2-large")
470WORSE_MODEL = ConceptEmbedder(dim=8, name="concept-v2-mini")
471
472
473def doc_text(d: Doc) -> str:
474    return f"{d.title}. {d.text}"
475
476
477# ---------------------------------------------------------------------------
478# 1. A minimal versioned vector index, and golden-set recall
479# ---------------------------------------------------------------------------
480
481
482@dataclass
483class VectorIndex:
484    """Exact (flat) index for one model version. Vectors are only comparable within it."""
485
486    model: ConceptEmbedder
487    ids: list[str] = field(default_factory=list)
488    vectors: list[np.ndarray] = field(default_factory=list)
489
490    @property
491    def version(self) -> str:
492        return self.model.name
493
494    def add(self, docs: list[Doc]) -> None:
495        for d, v in zip(docs, self.model.encode([doc_text(d) for d in docs])):
496            self.ids.append(d.id)
497            self.vectors.append(v)
498
499    def search_vector(self, q: np.ndarray, k: int) -> list[str]:
500        scores = np.stack(self.vectors) @ q
501        return [self.ids[i] for i in np.argsort(-scores, kind="stable")[:k]]
502
503    def search(self, query: str, k: int) -> list[str]:
504        return self.search_vector(self.model.encode(query), k)
505
506
507def golden_recall(query_model: ConceptEmbedder, doc_model: ConceptEmbedder, k: int, queries=LABELED_QUERIES) -> float:
508    """Share of golden queries with at least one relevant doc in the top k.
509
510    Documents are embedded with `doc_model` and queries with `query_model`.
511    Passing two different models reproduces the classic migration bug:
512    querying an old index with a new model.
513    """
514    idx = VectorIndex(doc_model)
515    idx.add(DOCS)
516    hits = [bool(rel & set(idx.search_vector(query_model.encode(q), k))) for q, rel in queries]
517    return float(np.mean(hits))
518
519
520# ---------------------------------------------------------------------------
521# 2. Blue/green migration to a new embedding model
522# ---------------------------------------------------------------------------
523
524
525class BlueGreenMigration:
526    """Move search from one embedding model to another without a bad day.
527
528    Steps: start (create the new index and begin dual-writing), backfill
529    (re-embed every existing document), shadow_compare (measure both on the
530    golden set), cutover (flip the live alias only if the new one is at least
531    as good), rollback (flip it back; the old index is kept until retired).
532    """
533
534    def __init__(self, current: ConceptEmbedder, docs: list[Doc] = DOCS):
535        self.docs = list(docs)
536        self.indexes: dict[str, VectorIndex] = {current.name: VectorIndex(current)}
537        self.indexes[current.name].add(self.docs)
538        self.live_version = current.name
539        self.previous_version: str | None = None
540        self.candidate: str | None = None
541        self.comparison: dict[str, float] = {}
542        self.log: list[str] = [f"live index {current.name} serving {len(self.docs)} docs"]
543
544    def start(self, new: ConceptEmbedder) -> None:
545        self.indexes[new.name] = VectorIndex(new)
546        self.candidate = new.name
547        self.comparison = {}
548        self.log.append(f"created empty index {new.name}; dual-writing new documents to both")
549
550    def add(self, doc: Doc) -> None:
551        """New documents go to every index, so the candidate never falls behind while backfilling."""
552        self.docs.append(doc)
553        for idx in self.indexes.values():
554            idx.add([doc])
555        self.log.append(f"dual-wrote {doc.id} to {sorted(self.indexes)}")
556
557    def backfill(self, batch_size: int = 8) -> None:
558        idx = self.indexes[self.candidate]
559        missing = [d for d in self.docs if d.id not in set(idx.ids)]
560        for start in range(0, len(missing), batch_size):  # batches: restartable, rate-limit friendly
561            idx.add(missing[start : start + batch_size])
562        self.log.append(f"backfilled {len(missing)} docs into {self.candidate}")
563
564    def shadow_compare(self, k: int = 3) -> dict[str, float]:
565        for name in (self.live_version, self.candidate):
566            m = self.indexes[name].model
567            self.comparison[name] = golden_recall(m, m, k)
568        self.log.append("shadow recall@%d: %s" % (k, ", ".join(f"{n} {r:.3f}" for n, r in self.comparison.items())))
569        return self.comparison
570
571    def cutover(self, min_gain: float = 0.0) -> bool:
572        if self.candidate not in self.comparison:
573            self.log.append("cutover refused: no shadow comparison yet")
574            return False
575        old, new = self.comparison[self.live_version], self.comparison[self.candidate]
576        if new < old + min_gain:
577            self.log.append(f"cutover refused: {self.candidate} recall {new:.3f} < live {old:.3f}")
578            return False
579        self.previous_version, self.live_version = self.live_version, self.candidate
580        self.log.append(f"cutover: live alias -> {self.live_version} (old index kept for rollback)")
581        return True
582
583    def rollback(self) -> None:
584        if self.previous_version:
585            self.live_version, self.previous_version = self.previous_version, self.live_version
586            self.log.append(f"rollback: live alias -> {self.live_version}")
587
588
589def reembed_estimate(n_docs: int, avg_tokens: int, tokens_per_second: float, dollars_per_million_tokens: float) -> dict[str, float]:
590    """Back-of-envelope time and money to re-embed a corpus. Prices vary; pass your own."""
591    tokens = n_docs * avg_tokens
592    return {"tokens": tokens, "hours": tokens / tokens_per_second / 3600, "dollars": tokens / 1e6 * dollars_per_million_tokens}
593
594
595# ---------------------------------------------------------------------------
596# 3. Domain mismatch
597# ---------------------------------------------------------------------------
598
599# Jargon a general model has never seen, and the concept each word means here.
600JARGON: dict[str, str] = {"ap": "invoice", "hotspot": "remote_access", "sso": "credential", "t-and-e": "money_back"}
601
602JARGON_QUERIES: list[tuple[str, set[str]]] = [
603    ("ap aging report", {"fin-005"}),  # accounts payable: vendor invoices
604    ("hotspot from a client site", {"it-003"}),  # remote access
605    ("sso lockout", {"it-001"}),  # single sign-on: your login
606    ("t-and-e submission deadline", {"fin-004"}),  # travel and expenses
607]
608
609
610class DomainTunedEmbedder(ConceptEmbedder):
611    """Stand-in for a model fine-tuned on company pairs: it has learned the jargon.
612
613    In real life this comes from contrastive fine-tuning on (query, document)
614    pairs from your own logs (`primer.ml.embeddings.contrastive`). Here we
615    simply teach the toy embedder what four jargon words mean.
616    """
617
618    def __init__(self, jargon: dict[str, str] = JARGON, **kw):
619        super().__init__(**{"name": "concept-v1-tuned", **kw})
620        self.jargon = jargon
621
622    def _token_vector(self, tok: str) -> np.ndarray:
623        concept = self.jargon.get(tok)
624        if concept is None:
625            return super()._token_vector(tok)
626        return _unit_vector(f"{self.name}/concept/{concept}", self.dim) + self.word_weight * _unit_vector(f"{self.name}/word/{tok}", self.dim)
627
628
629def build_eval_set(log: list[dict]) -> list[tuple[str, set[str]]]:
630    """Turn a search/support log into (query, relevant doc ids) cases.
631
632    Keep only sessions where the user's problem was resolved (the clicked
633    document really answered it); merge the same question asked twice.
634    """
635    cases: dict[str, set[str]] = {}
636    for row in log:
637        if not row["resolved"]:
638            continue
639        key = " ".join(tokenize(row["query"], drop_stopwords=False))  # "Travel policy?" == "travel policy"
640        cases.setdefault(key, set()).add(row["clicked"])
641    return list(cases.items())
642
643
644# ---------------------------------------------------------------------------
645# 4. Measure retrieval separately from generation
646# ---------------------------------------------------------------------------
647
648
649def triage(retrieved: list[str], relevant: set[str], answer_used: str) -> str:
650    """Which half of a RAG system failed?
651
652    If no relevant document was retrieved, no prompt can fix the answer: it's
653    a retrieval failure. If one was retrieved but the answer relied on a
654    different document, the generator misused good context.
655    """
656    if not relevant & set(retrieved):
657        return "retrieval failure"
658    if answer_used not in relevant:
659        return "generation failure"
660    return "correct"
661
662
663def triage_golden_set(model: ConceptEmbedder, k: int = 3, queries=LABELED_QUERIES) -> list[tuple[str, str]]:
664    """Run every golden query through retrieval and a toy generator that answers from the top document."""
665    idx = VectorIndex(model)
666    idx.add(DOCS)
667    out = []
668    for q, rel in queries:
669        top = idx.search(q, k)
670        out.append((q, triage(top, rel, answer_used=top[0])))
671    return out
672
673
674
675def rank_of_first_relevant(model: ConceptEmbedder, query: str, relevant: set[str]) -> int:
676    """1-based position of the first relevant document in the full ranking."""
677    idx = VectorIndex(model)
678    idx.add(DOCS)
679    ranking = idx.search(query, len(DOCS))
680    return next(i + 1 for i, d in enumerate(ranking) if d in relevant)
681
682
683# ---------------------------------------------------------------------------
684# 5. Figures
685# ---------------------------------------------------------------------------
686
687
688def figures() -> dict:
689    """Plots computed from this module's own functions. Keys match the docstring's image names."""
690    import matplotlib
691
692    matplotlib.use("Agg")
693    import matplotlib.pyplot as plt
694
695    figs = {}
696
697    # cross-model
698    labels = ["v1 docs,\nv1 queries", "v2 docs,\nv2 queries", "v1 docs,\nv2 queries"]
699    vals = [golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)]
700    fig, ax = plt.subplots(figsize=(6, 4))
701    bars = ax.bar(labels, vals, color=["C0", "C2", "C3"])
702    ax.bar_label(bars, fmt="%.2f")
703    ax.axhline(3 / len(DOCS), ls="--", color="0.5", label="random ranking (3 of 20)")
704    ax.set(ylabel="recall@3 on the golden set", ylim=(0, 1.1), title="Mixing two models' vectors fails silently")
705    ax.legend()
706    figs["cross_model"] = fig
707
708    # reembed
709    ns = np.logspace(5, 9, 30)
710    fig, (a1, a2) = plt.subplots(1, 2, figsize=(11, 4))
711    for r in (1e5, 1e6, 1e7):
712        a1.plot(ns, [reembed_estimate(int(n), 500, r, 0.02)["hours"] for n in ns], label=f"{r:,.0f} tokens/s")
713    for p in (0.01, 0.02, 0.13):
714        a2.plot(ns, [reembed_estimate(int(n), 500, 1e6, p)["dollars"] for n in ns], label=f"${p} per million tokens")
715    for ax, ylab in ((a1, "hours"), (a2, "dollars")):
716        ax.set_xscale("log")
717        ax.set_yscale("log")
718        ax.axvline(5e7, ls=":", color="0.5")
719        ax.set(xlabel="documents (500 tokens each, log scale)", ylabel=f"{ylab} (log scale)")
720        ax.legend(fontsize=8)
721    a1.set_title("Time to re-embed")
722    a2.set_title("Cost to re-embed")
723    figs["reembed"] = fig
724
725    # jargon
726    tuned = DomainTunedEmbedder()
727    qs = [q for q, _ in JARGON_QUERIES]
728    x = np.arange(len(qs))
729    fig, ax = plt.subplots(figsize=(8, 4))
730    ax.bar(x - 0.2, [rank_of_first_relevant(CURRENT_MODEL, q, r) for q, r in JARGON_QUERIES], 0.4, label="general model")
731    ax.bar(x + 0.2, [rank_of_first_relevant(tuned, q, r) for q, r in JARGON_QUERIES], 0.4, label="tuned on the jargon")
732    ax.set_xticks(x, qs, fontsize=8)
733    ax.set(ylabel="rank of the right document (1 is best)", title="Company jargon: general vs. tuned model")
734    ax.legend()
735    figs["jargon"] = fig
736
737    # triage
738    ks = [1, 3, 5]
739    colours = {"correct": "C2", "generation failure": "C1", "retrieval failure": "C3"}
740    fig, ax = plt.subplots(figsize=(6, 4))
741    bottoms = np.zeros(len(ks))
742    counts = {v: [] for v in colours}
743    for k in ks:
744        verdicts = [v for _, v in triage_golden_set(CURRENT_MODEL, k)]
745        for v in colours:
746            counts[v].append(verdicts.count(v))
747    for v, c in colours.items():
748        ax.bar([str(k) for k in ks], counts[v], bottom=bottoms, color=c, label=v)
749        bottoms += np.array(counts[v])
750    ax.set(xlabel="documents retrieved (k)", ylabel="golden questions", title="Which half failed? (generator uses the top document)")
751    ax.legend(fontsize=8)
752    figs["triage"] = fig
753
754    for f in figs.values():
755        f.tight_layout()
756    return figs
757
758
759# ---------------------------------------------------------------------------
760# 6. Narrated walkthrough
761# ---------------------------------------------------------------------------
762
763
764def demo() -> None:
765    banner("1. Every model has its own space")
766    a, b = ConceptEmbedder(name="concept-v1").encode("vpn"), ConceptEmbedder(name="concept-v2").encode("vpn")
767    say(f"Cosine between 'vpn' from v1 and 'vpn' from v2 (both 64 dims): {float(a @ b):.3f}.")
768    table(
769        ["documents", "queries", "recall@3"],
770        [
771            ("v1", "v1", golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3)),
772            ("v2", "v2", golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3)),
773            ("v1", "v2 (the bug)", golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)),
774        ],
775        floatfmt=".3f",
776    )
777    takeaway("Same number of dimensions, different space. Mixing them doesn't crash; it just quietly stops working.")
778
779    banner("2. A blue/green migration, step by step")
780    for candidate in (BETTER_MODEL, WORSE_MODEL):
781        m = BlueGreenMigration(CURRENT_MODEL)
782        m.start(candidate)
783        m.add(Doc("it-999", "Wi-Fi guest network", "The guest Wi-Fi password rotates every Monday.", "IT", "2026-09-01"))
784        m.backfill()
785        m.shadow_compare(k=3)
786        m.cutover()
787        print(f"  candidate {candidate.name}:")
788        for line in m.log:
789            print("   -", line)
790        print()
791    est = reembed_estimate(50_000_000, 500, 1_000_000, 0.02)
792    say(f"Re-embedding 50M docs × 500 tokens = {est['tokens']:,} tokens: {est['hours']:.2f} hours at 1M tokens/s, ${est['dollars']:,.0f} at $0.02 per million.")
793
794    banner("3. Domain jargon")
795    tuned = DomainTunedEmbedder()
796    table(
797        ["jargon question", "rank (general)", "rank (tuned)"],
798        [(q, rank_of_first_relevant(CURRENT_MODEL, q, r), rank_of_first_relevant(tuned, q, r)) for q, r in JARGON_QUERIES],
799    )
800    log = [
801        {"query": "VPN error 4012?", "clicked": "it-004", "resolved": True},
802        {"query": "vpn error 4012", "clicked": "it-003", "resolved": True},
803        {"query": "printer", "clicked": "it-009", "resolved": False},
804    ]
805    say(f"Eval set built from a 3-line log: {build_eval_set(log)}")
806
807    banner("4. Triage: retrieval failure or generation failure?")
808    table(["golden question", "verdict (k = 3)"], triage_golden_set(CURRENT_MODEL, 3))
809    takeaway("If the right document wasn't retrieved, no prompt can fix it. Measure retrieval on its own first.")
810
811
812if __name__ == "__main__":
813    demo()
Level 3: the code, function by function.
CURRENT_MODEL = <primer.common.embedder.ConceptEmbedder object>
SAME_SIZE_NEW_MODEL = <primer.common.embedder.ConceptEmbedder object>
BETTER_MODEL = <primer.common.embedder.ConceptEmbedder object>
def doc_text(d: primer.common.corpus.Doc) -> str: on GitHub
474def doc_text(d: Doc) -> str:
475    return f"{d.title}. {d.text}"
@dataclass
class VectorIndex: on GitHub
483@dataclass
484class VectorIndex:
485    """Exact (flat) index for one model version. Vectors are only comparable within it."""
486
487    model: ConceptEmbedder
488    ids: list[str] = field(default_factory=list)
489    vectors: list[np.ndarray] = field(default_factory=list)
490
491    @property
492    def version(self) -> str:
493        return self.model.name
494
495    def add(self, docs: list[Doc]) -> None:
496        for d, v in zip(docs, self.model.encode([doc_text(d) for d in docs])):
497            self.ids.append(d.id)
498            self.vectors.append(v)
499
500    def search_vector(self, q: np.ndarray, k: int) -> list[str]:
501        scores = np.stack(self.vectors) @ q
502        return [self.ids[i] for i in np.argsort(-scores, kind="stable")[:k]]
503
504    def search(self, query: str, k: int) -> list[str]:
505        return self.search_vector(self.model.encode(query), k)

Exact (flat) index for one model version. Vectors are only comparable within it.

VectorIndex( model: primer.common.embedder.ConceptEmbedder, ids: list[str] = <factory>, vectors: list[numpy.ndarray] = <factory>)
ids: list[str]
vectors: list[numpy.ndarray]
version: str on GitHub
491    @property
492    def version(self) -> str:
493        return self.model.name
def add(self, docs: list[primer.common.corpus.Doc]) -> None: on GitHub
495    def add(self, docs: list[Doc]) -> None:
496        for d, v in zip(docs, self.model.encode([doc_text(d) for d in docs])):
497            self.ids.append(d.id)
498            self.vectors.append(v)
def search_vector(self, q: numpy.ndarray, k: int) -> list[str]: on GitHub
500    def search_vector(self, q: np.ndarray, k: int) -> list[str]:
501        scores = np.stack(self.vectors) @ q
502        return [self.ids[i] for i in np.argsort(-scores, kind="stable")[:k]]
def search(self, query: str, k: int) -> list[str]: on GitHub
504    def search(self, query: str, k: int) -> list[str]:
505        return self.search_vector(self.model.encode(query), k)
def golden_recall( query_model: primer.common.embedder.ConceptEmbedder, doc_model: primer.common.embedder.ConceptEmbedder, k: int, queries=[("I forgot my password and I'm locked out", {'it-001'}), ('how long must a password be', {'it-002'}), ('what does ERR-4012 mean', {'it-004'}), ('ERR-4013 on my laptop', {'it-005'}), ('automobile reimbursement for business driving', {'fin-001'}), ('how many vacation days do I get', {'hr-001'}), ('connect to internal systems from home', {'it-003'}), ('suspicious email with a link', {'it-007'}), ('per-diem for meals when traveling', {'fin-002'}), ('enroll in MFA', {'it-006'}), ('first day checklist for a new hire', {'hr-003'}), ('match vendor bills to payments', {'fin-005'})]) -> float: on GitHub
508def golden_recall(query_model: ConceptEmbedder, doc_model: ConceptEmbedder, k: int, queries=LABELED_QUERIES) -> float:
509    """Share of golden queries with at least one relevant doc in the top k.
510
511    Documents are embedded with `doc_model` and queries with `query_model`.
512    Passing two different models reproduces the classic migration bug:
513    querying an old index with a new model.
514    """
515    idx = VectorIndex(doc_model)
516    idx.add(DOCS)
517    hits = [bool(rel & set(idx.search_vector(query_model.encode(q), k))) for q, rel in queries]
518    return float(np.mean(hits))

Share of golden queries with at least one relevant doc in the top k.

Documents are embedded with doc_model and queries with query_model. Passing two different models reproduces the classic migration bug: querying an old index with a new model.

class BlueGreenMigration: on GitHub
526class BlueGreenMigration:
527    """Move search from one embedding model to another without a bad day.
528
529    Steps: start (create the new index and begin dual-writing), backfill
530    (re-embed every existing document), shadow_compare (measure both on the
531    golden set), cutover (flip the live alias only if the new one is at least
532    as good), rollback (flip it back; the old index is kept until retired).
533    """
534
535    def __init__(self, current: ConceptEmbedder, docs: list[Doc] = DOCS):
536        self.docs = list(docs)
537        self.indexes: dict[str, VectorIndex] = {current.name: VectorIndex(current)}
538        self.indexes[current.name].add(self.docs)
539        self.live_version = current.name
540        self.previous_version: str | None = None
541        self.candidate: str | None = None
542        self.comparison: dict[str, float] = {}
543        self.log: list[str] = [f"live index {current.name} serving {len(self.docs)} docs"]
544
545    def start(self, new: ConceptEmbedder) -> None:
546        self.indexes[new.name] = VectorIndex(new)
547        self.candidate = new.name
548        self.comparison = {}
549        self.log.append(f"created empty index {new.name}; dual-writing new documents to both")
550
551    def add(self, doc: Doc) -> None:
552        """New documents go to every index, so the candidate never falls behind while backfilling."""
553        self.docs.append(doc)
554        for idx in self.indexes.values():
555            idx.add([doc])
556        self.log.append(f"dual-wrote {doc.id} to {sorted(self.indexes)}")
557
558    def backfill(self, batch_size: int = 8) -> None:
559        idx = self.indexes[self.candidate]
560        missing = [d for d in self.docs if d.id not in set(idx.ids)]
561        for start in range(0, len(missing), batch_size):  # batches: restartable, rate-limit friendly
562            idx.add(missing[start : start + batch_size])
563        self.log.append(f"backfilled {len(missing)} docs into {self.candidate}")
564
565    def shadow_compare(self, k: int = 3) -> dict[str, float]:
566        for name in (self.live_version, self.candidate):
567            m = self.indexes[name].model
568            self.comparison[name] = golden_recall(m, m, k)
569        self.log.append("shadow recall@%d: %s" % (k, ", ".join(f"{n} {r:.3f}" for n, r in self.comparison.items())))
570        return self.comparison
571
572    def cutover(self, min_gain: float = 0.0) -> bool:
573        if self.candidate not in self.comparison:
574            self.log.append("cutover refused: no shadow comparison yet")
575            return False
576        old, new = self.comparison[self.live_version], self.comparison[self.candidate]
577        if new < old + min_gain:
578            self.log.append(f"cutover refused: {self.candidate} recall {new:.3f} < live {old:.3f}")
579            return False
580        self.previous_version, self.live_version = self.live_version, self.candidate
581        self.log.append(f"cutover: live alias -> {self.live_version} (old index kept for rollback)")
582        return True
583
584    def rollback(self) -> None:
585        if self.previous_version:
586            self.live_version, self.previous_version = self.previous_version, self.live_version
587            self.log.append(f"rollback: live alias -> {self.live_version}")

Move search from one embedding model to another without a bad day.

Steps: start (create the new index and begin dual-writing), backfill (re-embed every existing document), shadow_compare (measure both on the golden set), cutover (flip the live alias only if the new one is at least as good), rollback (flip it back; the old index is kept until retired).

BlueGreenMigration( current: primer.common.embedder.ConceptEmbedder, docs: list[primer.common.corpus.Doc] = [Doc(id='it-001', title='How to reset your password', text="If you forgot your password or your account is locked, go to the self-service portal, choose 'Forgot password', verify with your authenticator app, and set a new password. The reset link expires after 15 minutes.", department='IT', updated='2026-03-02', acl=frozenset({'everyone'})), Doc(id='it-002', title='Password policy', text='Passwords must be at least 14 characters and include a number and a symbol. Passwords expire every 180 days and the last 10 passwords cannot be reused. This policy applies to all employees and contractors.', department='IT', updated='2025-11-20', acl=frozenset({'everyone'})), Doc(id='it-003', title='Setting up the VPN', text='Install the AnyConnect client from the software center, sign in with your company credentials, and approve the two-factor prompt. Use the VPN for all remote access to internal systems.', department='IT', updated='2026-01-15', acl=frozenset({'everyone'})), Doc(id='it-004', title='Error ERR-4012: VPN tunnel failed', text='ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. Update AnyConnect to version 5.1 or later and reboot.', department='IT', updated='2026-04-10', acl=frozenset({'everyone'})), Doc(id='it-005', title='Error ERR-4013: certificate expired', text="ERR-4013 means your device certificate expired. Open the software center and run 'Renew device certificate', then reconnect.", department='IT', updated='2026-04-10', acl=frozenset({'everyone'})), Doc(id='it-006', title='Setting up two-factor authentication', text='Download the authenticator app, scan the QR code on the security page, and enter the six digit code to finish enrolling in MFA. Two-factor is required for email and VPN.', department='IT', updated='2025-09-01', acl=frozenset({'everyone'})), Doc(id='it-007', title='Reporting phishing emails', text="If an email looks suspicious, do not click links. Use the 'Report phishing' button in Outlook. Security reviews every report within one business day.", department='IT', updated='2026-02-11', acl=frozenset({'everyone'})), Doc(id='it-008', title='Requesting a new laptop', text="Laptops are refreshed every three years. Submit a hardware request in the IT portal with your manager's approval. Standard devices are MacBook Pro or ThinkPad X1.", department='IT', updated='2025-12-05', acl=frozenset({'everyone'})), Doc(id='it-009', title='Printer troubleshooting', text='If printing fails, check the toner and paper tray, then remove and re-add the printer from settings. Floor printers are named by building and floor.', department='IT', updated='2024-06-30', acl=frozenset({'everyone'})), Doc(id='fin-001', title='Car mileage expense', text='Employees who use a personal vehicle for business driving are reimbursed at the standard mileage rate. Log each trip with date, distance and purpose, and submit the expense within 30 days.', department='Finance', updated='2026-01-05', acl=frozenset({'everyone'})), Doc(id='fin-002', title='Travel policy (2026)', text='Book flights and hotels through the travel portal. Economy class is required for flights under six hours. Meals are covered by a daily per-diem of 75 dollars.', department='Finance', updated='2026-01-01', acl=frozenset({'everyone'})), Doc(id='fin-003', title='Travel policy (2023, superseded)', text='Book flights through the travel agency by phone. Business class is allowed for flights over four hours. Meals are covered by a daily per-diem of 60 dollars.', department='Finance', updated='2023-01-01', acl=frozenset({'everyone'})), Doc(id='fin-004', title='Submitting expense receipts', text='Upload receipts to the expense tool within 30 days. Receipts are required for any expense over 25 dollars. Reimbursements are paid with the next payroll run.', department='Finance', updated='2025-10-12', acl=frozenset({'everyone'})), Doc(id='fin-005', title='Vendor invoice reconciliation', text='At quarter end, match each vendor invoice to its payment record. List every mismatch with the invoice number, amount and vendor, and send the reconciliation to the controller.', department='Finance', updated='2026-03-31', acl=frozenset({'finance'})), Doc(id='fin-006', title='Q3 revenue forecast', text='Q3 revenue is forecast at 41 million dollars, up 12 percent, driven by enterprise renewals.', department='Finance', updated='2026-07-01', acl=frozenset({'exec', 'finance'})), Doc(id='hr-001', title='Paid time off', text='Full-time employees accrue 20 days of PTO per year. Request vacation in the HR portal at least two weeks ahead. Unused PTO up to 5 days rolls over.', department='HR', updated='2026-01-01', acl=frozenset({'everyone'})), Doc(id='hr-002', title='Sick leave', text='Employees receive 10 paid sick days per year. No manager approval is needed for sick leave, but notify your team as early as possible.', department='HR', updated='2025-08-15', acl=frozenset({'everyone'})), Doc(id='hr-003', title='New hire onboarding', text='On your first day, collect your laptop from IT, complete security training, and set up two-factor authentication. Your manager will schedule orientation sessions for week one.', department='HR', updated='2026-02-01', acl=frozenset({'everyone'})), Doc(id='hr-004', title='Salary bands and bonus', text='Salary bands are reviewed every April. The annual bonus target is 10 percent of base pay, paid in March based on company and individual performance.', department='HR', updated='2026-04-01', acl=frozenset({'exec', 'hr'})), Doc(id='hr-005', title='Parental leave', text='Primary caregivers receive 16 weeks of paid parental leave; secondary caregivers receive 6 weeks. Leave can start up to two weeks before the expected birth or adoption date.', department='HR', updated='2025-05-20', acl=frozenset({'everyone'}))]) on GitHub
535    def __init__(self, current: ConceptEmbedder, docs: list[Doc] = DOCS):
536        self.docs = list(docs)
537        self.indexes: dict[str, VectorIndex] = {current.name: VectorIndex(current)}
538        self.indexes[current.name].add(self.docs)
539        self.live_version = current.name
540        self.previous_version: str | None = None
541        self.candidate: str | None = None
542        self.comparison: dict[str, float] = {}
543        self.log: list[str] = [f"live index {current.name} serving {len(self.docs)} docs"]
docs
indexes: dict[str, VectorIndex]
previous_version: str | None
candidate: str | None
comparison: dict[str, float]
log: list[str]
def start(self, new: primer.common.embedder.ConceptEmbedder) -> None: on GitHub
545    def start(self, new: ConceptEmbedder) -> None:
546        self.indexes[new.name] = VectorIndex(new)
547        self.candidate = new.name
548        self.comparison = {}
549        self.log.append(f"created empty index {new.name}; dual-writing new documents to both")
def add(self, doc: primer.common.corpus.Doc) -> None: on GitHub
551    def add(self, doc: Doc) -> None:
552        """New documents go to every index, so the candidate never falls behind while backfilling."""
553        self.docs.append(doc)
554        for idx in self.indexes.values():
555            idx.add([doc])
556        self.log.append(f"dual-wrote {doc.id} to {sorted(self.indexes)}")

New documents go to every index, so the candidate never falls behind while backfilling.

def backfill(self, batch_size: int = 8) -> None: on GitHub
558    def backfill(self, batch_size: int = 8) -> None:
559        idx = self.indexes[self.candidate]
560        missing = [d for d in self.docs if d.id not in set(idx.ids)]
561        for start in range(0, len(missing), batch_size):  # batches: restartable, rate-limit friendly
562            idx.add(missing[start : start + batch_size])
563        self.log.append(f"backfilled {len(missing)} docs into {self.candidate}")
def shadow_compare(self, k: int = 3) -> dict[str, float]: on GitHub
565    def shadow_compare(self, k: int = 3) -> dict[str, float]:
566        for name in (self.live_version, self.candidate):
567            m = self.indexes[name].model
568            self.comparison[name] = golden_recall(m, m, k)
569        self.log.append("shadow recall@%d: %s" % (k, ", ".join(f"{n} {r:.3f}" for n, r in self.comparison.items())))
570        return self.comparison
def cutover(self, min_gain: float = 0.0) -> bool: on GitHub
572    def cutover(self, min_gain: float = 0.0) -> bool:
573        if self.candidate not in self.comparison:
574            self.log.append("cutover refused: no shadow comparison yet")
575            return False
576        old, new = self.comparison[self.live_version], self.comparison[self.candidate]
577        if new < old + min_gain:
578            self.log.append(f"cutover refused: {self.candidate} recall {new:.3f} < live {old:.3f}")
579            return False
580        self.previous_version, self.live_version = self.live_version, self.candidate
581        self.log.append(f"cutover: live alias -> {self.live_version} (old index kept for rollback)")
582        return True
def rollback(self) -> None: on GitHub
584    def rollback(self) -> None:
585        if self.previous_version:
586            self.live_version, self.previous_version = self.previous_version, self.live_version
587            self.log.append(f"rollback: live alias -> {self.live_version}")
def reembed_estimate( n_docs: int, avg_tokens: int, tokens_per_second: float, dollars_per_million_tokens: float) -> dict[str, float]: on GitHub
590def reembed_estimate(n_docs: int, avg_tokens: int, tokens_per_second: float, dollars_per_million_tokens: float) -> dict[str, float]:
591    """Back-of-envelope time and money to re-embed a corpus. Prices vary; pass your own."""
592    tokens = n_docs * avg_tokens
593    return {"tokens": tokens, "hours": tokens / tokens_per_second / 3600, "dollars": tokens / 1e6 * dollars_per_million_tokens}

Back-of-envelope time and money to re-embed a corpus. Prices vary; pass your own.

JARGON: dict[str, str] = {'ap': 'invoice', 'hotspot': 'remote_access', 'sso': 'credential', 't-and-e': 'money_back'}
JARGON_QUERIES: list[tuple[str, set[str]]] = [('ap aging report', {'fin-005'}), ('hotspot from a client site', {'it-003'}), ('sso lockout', {'it-001'}), ('t-and-e submission deadline', {'fin-004'})]
611class DomainTunedEmbedder(ConceptEmbedder):
612    """Stand-in for a model fine-tuned on company pairs: it has learned the jargon.
613
614    In real life this comes from contrastive fine-tuning on (query, document)
615    pairs from your own logs (`primer.ml.embeddings.contrastive`). Here we
616    simply teach the toy embedder what four jargon words mean.
617    """
618
619    def __init__(self, jargon: dict[str, str] = JARGON, **kw):
620        super().__init__(**{"name": "concept-v1-tuned", **kw})
621        self.jargon = jargon
622
623    def _token_vector(self, tok: str) -> np.ndarray:
624        concept = self.jargon.get(tok)
625        if concept is None:
626            return super()._token_vector(tok)
627        return _unit_vector(f"{self.name}/concept/{concept}", self.dim) + self.word_weight * _unit_vector(f"{self.name}/word/{tok}", self.dim)

Stand-in for a model fine-tuned on company pairs: it has learned the jargon.

In real life this comes from contrastive fine-tuning on (query, document) pairs from your own logs (primer.ml.embeddings.contrastive). Here we simply teach the toy embedder what four jargon words mean.

DomainTunedEmbedder( jargon: dict[str, str] = {'ap': 'invoice', 'hotspot': 'remote_access', 'sso': 'credential', 't-and-e': 'money_back'}, **kw) on GitHub
619    def __init__(self, jargon: dict[str, str] = JARGON, **kw):
620        super().__init__(**{"name": "concept-v1-tuned", **kw})
621        self.jargon = jargon
jargon
def build_eval_set(log: list[dict]) -> list[tuple[str, set[str]]]: on GitHub
630def build_eval_set(log: list[dict]) -> list[tuple[str, set[str]]]:
631    """Turn a search/support log into (query, relevant doc ids) cases.
632
633    Keep only sessions where the user's problem was resolved (the clicked
634    document really answered it); merge the same question asked twice.
635    """
636    cases: dict[str, set[str]] = {}
637    for row in log:
638        if not row["resolved"]:
639            continue
640        key = " ".join(tokenize(row["query"], drop_stopwords=False))  # "Travel policy?" == "travel policy"
641        cases.setdefault(key, set()).add(row["clicked"])
642    return list(cases.items())

Turn a search/support log into (query, relevant doc ids) cases.

Keep only sessions where the user's problem was resolved (the clicked document really answered it); merge the same question asked twice.

def triage(retrieved: list[str], relevant: set[str], answer_used: str) -> str: on GitHub
650def triage(retrieved: list[str], relevant: set[str], answer_used: str) -> str:
651    """Which half of a RAG system failed?
652
653    If no relevant document was retrieved, no prompt can fix the answer: it's
654    a retrieval failure. If one was retrieved but the answer relied on a
655    different document, the generator misused good context.
656    """
657    if not relevant & set(retrieved):
658        return "retrieval failure"
659    if answer_used not in relevant:
660        return "generation failure"
661    return "correct"

Which half of a RAG system failed?

If no relevant document was retrieved, no prompt can fix the answer: it's a retrieval failure. If one was retrieved but the answer relied on a different document, the generator misused good context.

def triage_golden_set( model: primer.common.embedder.ConceptEmbedder, k: int = 3, queries=[("I forgot my password and I'm locked out", {'it-001'}), ('how long must a password be', {'it-002'}), ('what does ERR-4012 mean', {'it-004'}), ('ERR-4013 on my laptop', {'it-005'}), ('automobile reimbursement for business driving', {'fin-001'}), ('how many vacation days do I get', {'hr-001'}), ('connect to internal systems from home', {'it-003'}), ('suspicious email with a link', {'it-007'}), ('per-diem for meals when traveling', {'fin-002'}), ('enroll in MFA', {'it-006'}), ('first day checklist for a new hire', {'hr-003'}), ('match vendor bills to payments', {'fin-005'})]) -> list[tuple[str, str]]: on GitHub
664def triage_golden_set(model: ConceptEmbedder, k: int = 3, queries=LABELED_QUERIES) -> list[tuple[str, str]]:
665    """Run every golden query through retrieval and a toy generator that answers from the top document."""
666    idx = VectorIndex(model)
667    idx.add(DOCS)
668    out = []
669    for q, rel in queries:
670        top = idx.search(q, k)
671        out.append((q, triage(top, rel, answer_used=top[0])))
672    return out

Run every golden query through retrieval and a toy generator that answers from the top document.

def rank_of_first_relevant( model: primer.common.embedder.ConceptEmbedder, query: str, relevant: set[str]) -> int: on GitHub
676def rank_of_first_relevant(model: ConceptEmbedder, query: str, relevant: set[str]) -> int:
677    """1-based position of the first relevant document in the full ranking."""
678    idx = VectorIndex(model)
679    idx.add(DOCS)
680    ranking = idx.search(query, len(DOCS))
681    return next(i + 1 for i, d in enumerate(ranking) if d in relevant)

1-based position of the first relevant document in the full ranking.

def figures() -> dict: on GitHub
689def figures() -> dict:
690    """Plots computed from this module's own functions. Keys match the docstring's image names."""
691    import matplotlib
692
693    matplotlib.use("Agg")
694    import matplotlib.pyplot as plt
695
696    figs = {}
697
698    # cross-model
699    labels = ["v1 docs,\nv1 queries", "v2 docs,\nv2 queries", "v1 docs,\nv2 queries"]
700    vals = [golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3), golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)]
701    fig, ax = plt.subplots(figsize=(6, 4))
702    bars = ax.bar(labels, vals, color=["C0", "C2", "C3"])
703    ax.bar_label(bars, fmt="%.2f")
704    ax.axhline(3 / len(DOCS), ls="--", color="0.5", label="random ranking (3 of 20)")
705    ax.set(ylabel="recall@3 on the golden set", ylim=(0, 1.1), title="Mixing two models' vectors fails silently")
706    ax.legend()
707    figs["cross_model"] = fig
708
709    # reembed
710    ns = np.logspace(5, 9, 30)
711    fig, (a1, a2) = plt.subplots(1, 2, figsize=(11, 4))
712    for r in (1e5, 1e6, 1e7):
713        a1.plot(ns, [reembed_estimate(int(n), 500, r, 0.02)["hours"] for n in ns], label=f"{r:,.0f} tokens/s")
714    for p in (0.01, 0.02, 0.13):
715        a2.plot(ns, [reembed_estimate(int(n), 500, 1e6, p)["dollars"] for n in ns], label=f"${p} per million tokens")
716    for ax, ylab in ((a1, "hours"), (a2, "dollars")):
717        ax.set_xscale("log")
718        ax.set_yscale("log")
719        ax.axvline(5e7, ls=":", color="0.5")
720        ax.set(xlabel="documents (500 tokens each, log scale)", ylabel=f"{ylab} (log scale)")
721        ax.legend(fontsize=8)
722    a1.set_title("Time to re-embed")
723    a2.set_title("Cost to re-embed")
724    figs["reembed"] = fig
725
726    # jargon
727    tuned = DomainTunedEmbedder()
728    qs = [q for q, _ in JARGON_QUERIES]
729    x = np.arange(len(qs))
730    fig, ax = plt.subplots(figsize=(8, 4))
731    ax.bar(x - 0.2, [rank_of_first_relevant(CURRENT_MODEL, q, r) for q, r in JARGON_QUERIES], 0.4, label="general model")
732    ax.bar(x + 0.2, [rank_of_first_relevant(tuned, q, r) for q, r in JARGON_QUERIES], 0.4, label="tuned on the jargon")
733    ax.set_xticks(x, qs, fontsize=8)
734    ax.set(ylabel="rank of the right document (1 is best)", title="Company jargon: general vs. tuned model")
735    ax.legend()
736    figs["jargon"] = fig
737
738    # triage
739    ks = [1, 3, 5]
740    colours = {"correct": "C2", "generation failure": "C1", "retrieval failure": "C3"}
741    fig, ax = plt.subplots(figsize=(6, 4))
742    bottoms = np.zeros(len(ks))
743    counts = {v: [] for v in colours}
744    for k in ks:
745        verdicts = [v for _, v in triage_golden_set(CURRENT_MODEL, k)]
746        for v in colours:
747            counts[v].append(verdicts.count(v))
748    for v, c in colours.items():
749        ax.bar([str(k) for k in ks], counts[v], bottom=bottoms, color=c, label=v)
750        bottoms += np.array(counts[v])
751    ax.set(xlabel="documents retrieved (k)", ylabel="golden questions", title="Which half failed? (generator uses the top document)")
752    ax.legend(fontsize=8)
753    figs["triage"] = fig
754
755    for f in figs.values():
756        f.tight_layout()
757    return figs

Plots computed from this module's own functions. Keys match the docstring's image names.

def demo() -> None: on GitHub
765def demo() -> None:
766    banner("1. Every model has its own space")
767    a, b = ConceptEmbedder(name="concept-v1").encode("vpn"), ConceptEmbedder(name="concept-v2").encode("vpn")
768    say(f"Cosine between 'vpn' from v1 and 'vpn' from v2 (both 64 dims): {float(a @ b):.3f}.")
769    table(
770        ["documents", "queries", "recall@3"],
771        [
772            ("v1", "v1", golden_recall(CURRENT_MODEL, CURRENT_MODEL, 3)),
773            ("v2", "v2", golden_recall(SAME_SIZE_NEW_MODEL, SAME_SIZE_NEW_MODEL, 3)),
774            ("v1", "v2 (the bug)", golden_recall(SAME_SIZE_NEW_MODEL, CURRENT_MODEL, 3)),
775        ],
776        floatfmt=".3f",
777    )
778    takeaway("Same number of dimensions, different space. Mixing them doesn't crash; it just quietly stops working.")
779
780    banner("2. A blue/green migration, step by step")
781    for candidate in (BETTER_MODEL, WORSE_MODEL):
782        m = BlueGreenMigration(CURRENT_MODEL)
783        m.start(candidate)
784        m.add(Doc("it-999", "Wi-Fi guest network", "The guest Wi-Fi password rotates every Monday.", "IT", "2026-09-01"))
785        m.backfill()
786        m.shadow_compare(k=3)
787        m.cutover()
788        print(f"  candidate {candidate.name}:")
789        for line in m.log:
790            print("   -", line)
791        print()
792    est = reembed_estimate(50_000_000, 500, 1_000_000, 0.02)
793    say(f"Re-embedding 50M docs × 500 tokens = {est['tokens']:,} tokens: {est['hours']:.2f} hours at 1M tokens/s, ${est['dollars']:,.0f} at $0.02 per million.")
794
795    banner("3. Domain jargon")
796    tuned = DomainTunedEmbedder()
797    table(
798        ["jargon question", "rank (general)", "rank (tuned)"],
799        [(q, rank_of_first_relevant(CURRENT_MODEL, q, r), rank_of_first_relevant(tuned, q, r)) for q, r in JARGON_QUERIES],
800    )
801    log = [
802        {"query": "VPN error 4012?", "clicked": "it-004", "resolved": True},
803        {"query": "vpn error 4012", "clicked": "it-003", "resolved": True},
804        {"query": "printer", "clicked": "it-009", "resolved": False},
805    ]
806    say(f"Eval set built from a 3-line log: {build_eval_set(log)}")
807
808    banner("4. Triage: retrieval failure or generation failure?")
809    table(["golden question", "verdict (k = 3)"], triage_golden_set(CURRENT_MODEL, 3))
810    takeaway("If the right document wasn't retrieved, no prompt can fix it. Measure retrieval on its own first.")