primer.agents.cost

Cost and latency: paying less per successful task without getting worse

Run: python -m primer.agents.cost

This lesson builds on tokens and pricing from primer.ml.inference, on the context window from primer.agents.context and on the release gate from primer.agents.evals.

Level 1: The practitioner's guide

In one sentence. Cost engineering for a model-backed system means paying less per successful task, not per call, by taking the levers that can't hurt quality first and checking the ones that can against an eval.

When you need it. When the bill grows faster than the value, when a latency budget is missed, or before the first big customer arrives and the price per task stops being a rounding error. The tell: you know your spend per month but not your cost per successful task, or you know cost per call but have never counted retries and human clean-up. The lesson's worked example shows why that number, and only that number, decides things: a small model at \$0.002 a call with a 60% success rate beats a large one at \$0.010 and 95% if failures can simply be retried (\$0.0033 against \$0.0105 per success), and loses by 7x if a person has to fix each failure (\$0.802 against \$0.110 per task, at an illustrative \$2 per fix). You don't need any of this for a prototype with ten users a day, and you shouldn't touch the levers that change behaviour (routing, semantic caching) until you have an eval to watch them with. All prices in this lesson are illustrative constants chosen for round arithmetic (a large model at \$5 per million input tokens and \$25 per million output, a small one at \$1 and \$5); check a provider's current price list for real numbers.

Your options. Eight levers, from the ones that can't hurt quality to the ones that can:

Option What it does What it guarantees What it costs Where it lives
Prompt caching Puts the stable part of the prompt first so the provider reuses its work No behaviour change; the worked 10,000-token call falls from \$0.0625 to \$0.0265 with 8,000 tokens cached A small write premium the first time; reordering the prompt The provider's API
Trim tokens Compress tool results, send fewer retrieved chunks, ask for concise output No change if you only cut what the next step never read; output is the expensive direction Engineering, and care not to cut what was needed Your prompt and code
Parallel tool calls Run independent tool calls at the same time Wall-clock time falls to the slowest call (300 ms to 100 ms for three lookups); tokens unchanged Concurrency in the agent loop Your agent loop
Exact response cache The same question, ignoring case and spaces, returns the stored answer Never wrong until the facts change A time-to-live to manage Your code
Batch API Non-interactive work submitted in bulk at about half price Half price, results within hours Waiting; only for work nobody is waiting on The provider's API
Budgets and alerts Caps steps and tokens per task and spend per tenant; alerts on a sudden jump A confused agent stops instead of looping all night Choosing the limits; a false alarm now and then Your code
Routing Sends easy tasks (classify, extract, format) to a small model and hard ones to a large one 53% saved on this lesson's ten-task workload, if the router is right The first lever that can lose quality; needs an eval before and after Your code, in front of every request
Semantic cache Answers a similar question from memory, using embeddings to judge similarity The cheapest possible hit Confidently wrong answers: on this lesson's pairs, 40% of different-intent questions are served from cache at a 0.95 threshold Your code, plus an embedding model

A ninth, for later: train or distil a small model on the large model's answers so it can take more of the routed traffic (Hinton et al., 2015).

How to choose. Measure first, then go down the table in order.

  • No measurement yet: instrument cost per task by step, by model, input against output and cached against uncached, and put an eval set in place to hold quality fixed.
  • A long, stable system prompt or tool list: prompt caching, today. It is the only lever that halves a bill without changing a single answer.
  • Big tool results or many retrieved chunks: trim. Compress to the fields the next step needs and send the top few reranked chunks.
  • Several independent lookups per turn: run them concurrently, and stream the answer so users see progress.
  • Anything nobody is waiting on (nightly evals, backfills, bulk classification): batch it.
  • Most traffic is simple: route, with the eval watching. A router that sends hard tasks to the small model saves money and quietly loses quality.
  • Narrow, curated, FAQ-style traffic and nothing else: a semantic cache with guards (identifiers, numbers and negation must match), a time-to-live and a measured wrong-hit rate. On everyday questions no threshold makes it safe.
  • Whatever you pick, judge it by cost per successful task, with retries and human clean-up counted in.

What it costs. The levers stack. On this lesson's ten-task workload the baseline is \$0.671; caching takes it to \$0.311 (2.2x), trimming tool output to \$0.186 (3.6x), concise output to \$0.171 (3.9x), routing to \$0.086 (7.8x) and batching the non-interactive share to \$0.080 (8.4x). The first three are free wins that don't touch quality; routing is the big one and the first that can. Caching is not free on the first request: Anthropic's prompt caching documentation, for one, prices a cache write at 1.25 times the base input price for a five-minute cache (2 times for an hour), reads at a tenth of it or less, and only caches prompts above a per-model minimum of a few hundred to a few thousand tokens. Batching costs time: the same provider quotes a 50% discount with most batches finishing within an hour. Parallel calls and streaming cost no tokens at all; they buy time and perceived speed. Budgets cost the occasional false alarm: this lesson's anomaly rule flags a task using more than the mean plus three standard deviations of recent usage, so a 10,000-token task against a history around 1,000 trips it at once.

What breaks.

  • Cost per call replacing cost per success. The cheap model looks cheaper until failures are priced. Count retries and fixes.
  • Routing that loses quality quietly. Nothing errors; the answers just get worse. Gate the router with the eval, and re-check when the small model changes.
  • The semantic cache that answers the wrong question. "How many sick days do I get?" scores 0.97 against the vacation question and gets the vacation answer. Near-misses score higher than real paraphrases, so no threshold separates them.
  • Stale cache hits. The policy changed; the cache didn't. Give every entry a time-to-live.
  • A cache that never hits. Something volatile (a timestamp, the user's name) sits at the top of the prompt, so the prefix differs every call. Stable content first.
  • Trimming what was needed. A tool result cut to 500 tokens loses the field the next step reads. Compress by field, not by length alone.
  • The runaway loop. An agent calls the same tool with the same arguments all night. Cap steps and tokens per task; alert on tokens per task jumping.
  • One tenant's surprise invoice. A shared platform with no per-customer cap turns one customer's bug into everyone's bill.

In the wild. Prompt caching and batch processing are standard features of hosted APIs; Anthropic's documentation for both is linked in Further reading and gives the multipliers quoted above. FrugalGPT (Chen, Zaharia and Zou, 2023) named the three families, prompt adaptation, model approximation and cascades that try a cheap model first and escalate, and reported matching the best single model with up to 98% less cost on its benchmarks. Distillation (Hinton, Vinyals and Dean, 2015) is how the small model in the routing table gets good enough to take more traffic. Tooling: LiteLLM is an open-source gateway that puts one interface in front of many providers and adds routing with fallbacks, per-key and per-team budgets, spend tracking and caching; GPTCache is an open-source semantic cache built on embeddings and a vector store with a pluggable similarity evaluator, the design this lesson's SemanticCache reproduces in miniature, guards and all. Parallel tool use is part of the tool-calling protocol of the major APIs: the model asks for several tools in one turn and you return all the results in one message.

Go deeper. Level 2 prices a request symbol by symbol, builds the router, both caches and the parallel loop in plain Python, derives the anomaly alert from a mean and a standard deviation, works the unit economics with every number, and stacks the levers to draw the 8.4x figure. If you only needed to know which lever to pull first, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

A restaurant doesn't cut costs by buying worse ingredients for every dish. It sends simple orders to the line cook and complex ones to the head chef, pre-chops what every dish shares, doesn't plate food nobody eats, cooks things in parallel, does bulk prep overnight when it's cheaper, and watches the one kitchen station whose bills suddenly spike.

Every lever in this lesson is one of those moves. The measure that matters is cost per successful task, not cost per call: a cheap attempt that fails still has to be paid for, and so does fixing it.

All prices in this module are illustrative constants chosen for round arithmetic (a "large" model at \$5 per million input tokens and \$25 per million output tokens; a "small" one at \$1 and \$5). They show the shape of the trade-offs; check your provider's current price list for real numbers.

How a request is priced

Everyday picture. A taxi that charges one rate for the distance to your pickup and a higher rate for the ride itself. Models charge per token (a word piece, about 4 characters of English): one rate for tokens you send (input) and a higher rate for tokens the model writes (output), because writing happens one token at a time (see primer.ml.inference).

Worked example. A large-model call with 10,000 input tokens and 500 output tokens: 10,000 × \$5 / 1,000,000 = \$0.05 for input, plus 500 × \$25 / 1,000,000 = \$0.0125 for output, total \$0.0625. If 8,000 of those input tokens are a stable prefix served from the prompt cache (reused work from an earlier identical beginning, billed here at 10% of the input price), input becomes 8,000 × \$0.50/M + 2,000 × \$5/M = \$0.014, and the call costs \$0.0265, 58% less.

Level 3: the formula and its symbols

$$ \text{cost} = \frac{c \cdot \rho\, p_{\text{in}} + (n_{\text{in}} - c)\, p_{\text{in}} + n_{\text{out}}\, p_{\text{out}}}{10^6} $$

Symbols

Symbol Meaning here Worked example
$n_{\text{in}}$ input tokens sent 10,000
$c$ input tokens served from the prompt cache 8,000
$n_{\text{out}}$ output tokens generated 500
$p_{\text{in}}, p_{\text{out}}$ price per million input / output tokens \$5, \$25
$\rho$ cached-read price as a fraction of normal input price (Greek rho) 0.1
$10^6$ prices are per million tokens

In words: cached input at its discounted rate, plus the rest of the input at full rate, plus output at the output rate, all per million.

On the worked example: (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ = (4,000 + 10,000 + 12,500) / 10⁶ = \$0.0265.

Level 3: in Python

In Python:

n_in, c, n_out = 10_000, 8_000, 500
p_in, p_out, rho = 5, 25, 0.1
cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6
round(cost, 4)  # → 0.0265

Two facts fall out: output tokens cost several times more than input (ask for concise answers), and a cached prefix is nearly free (put stable content first; see primer.agents.context). Caches usually charge a small premium the first time a prefix is written; this module ignores it for simplicity.

In code: request_cost is the formula above, reading each model's Price from PRICES.

Route each task to the cheapest model that can do it

Everyday picture. A hospital triage nurse: sprained ankles go to the nurse practitioner, chest pains to the cardiologist. Nobody sends every patient to the most expensive specialist.

Worked example. The sample workload is 10 tasks: 7 simple (classify a ticket, extract a date) and 3 complex (plan a migration). All on the large model it costs \$0.671. Routing the 7 simple ones to the small model, which is 5x cheaper per token, brings it to \$0.314: 53% saved, with the hard tasks still on the strong model.

flowchart LR T[Incoming task] --> R{Router<br/>rules or a small classifier} R -->|classify, extract,<br/>format, short| S[Small, fast model] R -->|plan, analyze,<br/>multi-step reasoning| L[Large model] S --> O[Result] L --> O

Reading it: the router is cheap (a few rules here; often a small classifier) and sits in front of every request. The whole saving comes from the left branch: most traffic in real systems is simple, and simple work runs well on small models. Validate the router with evals (primer.agents.evals): a router that sends hard tasks to the small model saves money and quietly loses quality.

In code: route is the rule-based router; workload_cost prices a list of WorkItem tasks with any levers switched on; routing_savings compares all-large against routed on SAMPLE_WORKLOAD.

Response caching and semantic caching

Everyday picture. A receptionist who's been asked "what's the wifi password?" a hundred times just answers from memory. That's an exact response cache: the same question (ignoring case and spaces) returns the stored answer with no model call. A semantic cache goes further and answers similar questions from memory, using embeddings (vectors where closeness means similar meaning, see primer.ml.embeddings) to decide "similar". That's where the receptionist starts giving the vacation-policy answer to someone asking about sick leave.

Worked example. With the toy embedder in primer.common.embedder:

Stored question New question Similarity Same intent?
How do I reset my password? I forgot my password, how do I recover it? 0.88 yes
How many vacation days do I get? How many sick days do I get? 0.97 no
Request a new laptop Request a new monitor 0.96 no
What does ERR-4012 mean? What does ERR-4013 mean? 0.56 no

The two near-misses score higher than a genuine paraphrase. No single threshold separates them. Two cheap guards help: identifiers and numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match ("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which share a topic and differ in a single word.

flowchart TD Q[New question] --> E{Exact match<br/>in response cache?} E -->|yes| A1[Return stored answer] E -->|no| V[Embed and find the<br/>nearest stored question] V --> T{Similarity above<br/>threshold?} T -->|no| M[Call the model] T -->|yes| G{IDs, numbers and<br/>negation identical?<br/>entry not expired?} G -->|no| M G -->|yes| A2[Return stored answer] M --> S[Store the new answer]

Reading it: the exact cache is checked first because it's free and never wrong. The semantic path adds two gates: the similarity threshold, and the guards. An expired entry (older than its TTL, time to live) is treated as a miss, so answers about changing facts don't go stale forever.

No threshold separates them: wrong hits stay at 40 to 60% through 0.95, often above the paraphrase hit rate, and reach 0 only where paraphrase hits do too

Reading it: the x-axis is the similarity threshold; the blue line is the share of genuine paraphrases answered from cache (good), and the red line is the share of different-intent questions answered from cache (a wrong answer served confidently). Lowering the threshold raises both, but look where the lines sit: from 0.4 upward the red line is above the blue one, so the cache serves more wrong answers than right ones (60% of different-intent questions against at most 57% of paraphrases). Even at 0.95 it still answers 40% of them wrongly and only 14% of paraphrases, because "sick" vs "vacation" scores 0.97. On everyday questions no threshold makes semantic caching safe; it's safe only for narrow, curated FAQ-style traffic, with guards, a TTL, and a measured wrong-hit rate.

In code: ResponseCache is the exact cache. SemanticCache is the semantic path: SemanticCache.lookup skips expired entries and entries whose identifiers or negation differ, then returns the nearest survivor, and SemanticCache.get applies the threshold. semantic_cache_sweep draws the figure from the labelled pairs in CACHE_PAIRS.

Trim tokens

Everyday picture. Don't photocopy the whole binder when the colleague needs one page. Compress tool results to the fields the next step needs (primer.agents.context.compress_tool_output), send the top few reranked chunks instead of dozens, remove repeated boilerplate from system prompts, and ask for concise output, because output is the expensive direction.

In code: workload_cost models trimming as cutting each task's tool output to 500 tokens, and concise output as cutting long answers by a third.

Run independent tool calls at the same time

Everyday picture. Boil the pasta while the sauce simmers. Cooking them one after the other takes the sum of the times; together, the time of the slowest.

Worked example. Three independent lookups of 100 ms each: sequentially about 300 ms; concurrently about 100 ms.

sequenceDiagram participant A as Agent participant T1 as Weather API participant T2 as Calendar API participant T3 as CRM API Note over A,T3: Sequential: about 300 ms A->>T1: call T1-->>A: result A->>T2: call T2-->>A: result A->>T3: call T3-->>A: result Note over A,T3: Parallel: about 100 ms par A->>T1: call and A->>T2: call and A->>T3: call end T1-->>A: result T2-->>A: result T3-->>A: result

Reading it: in the top half each call waits for the previous result; in the bottom half the three requests leave together and the agent waits only for the slowest. Models can ask for several tools in one turn (parallel tool use); run them concurrently and return all results in one message. Parallel calls cut wall-clock time, not tokens.

Three 100 ms tool calls take 300 ms end to end when run one after another, but only 100 ms when started together

Reading it: each bar is one 100 ms tool call on a shared time axis. The sequential calls stack end to end, 300 ms in all; the parallel ones start together, so the whole batch finishes in the time of one. Run the lesson to measure it with asyncio: the real timings land within a few milliseconds of these bars.

In code: run_sequential awaits each call before starting the next; run_parallel starts them all with Python's asyncio gather and waits once.

Streaming (showing tokens as they're generated) doesn't reduce total time either, but users see progress immediately, which changes how fast the system feels.

Batch what isn't interactive

Everyday picture. Sending the laundry out to be done overnight at half price instead of waiting at the express counter. Providers offer batch APIs: submit many requests, get results within hours (often within 24), typically at about half the price. Use them for anything no one is waiting on: nightly evals, backfilling document processing, bulk classification.

In code: batch_cost prices a list of requests at the batch discount.

Budgets and alerts

Everyday picture. A prepaid card with a hard limit, and a bank that texts you when a purchase looks nothing like your usual spending.

A task budget caps steps and tokens per task, so a confused agent stops instead of looping all night. A tenant spend cap (a tenant is one customer organisation on a shared platform) stops one customer's runaway usage from becoming a surprise invoice. An anomaly alert fires when a task uses far more tokens than usual:

Level 3: the formula and its symbols

$$ \text{alert if } x > \mu + z\,\sigma $$

Symbols

Symbol Meaning here Worked example
$x$ tokens used by the task just finished 10,000
$\mu$ mean tokens per task over this tenant's recent history (Greek mu) 1,000
$\sigma$ standard deviation of that history: the typical distance from the mean (Greek sigma) ≈ 71
$z$ how many standard deviations count as unusual 3

In words: alert when a task uses more tokens than the usual amount plus three times the usual spread.

On the worked example: history 1000, 1100, 900, 1050, 950, 1000 has mean 1,000 and standard deviation ≈ 71 (the squared distances from the mean add up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7), so the line is 1,000 + 3 × 70.7 ≈ 1,212. A 10,000-token task is far above it: alert. Such jumps usually mean a loop or a bad deploy.

Level 3: in Python

In Python:

import statistics
history = [1000, 1100, 900, 1050, 950, 1000]
mu = statistics.mean(history)
# divides by 6 - 1, as above
sigma = statistics.stdev(history)
mu, round(sigma, 1)  # → (1000, 70.7)
z = 3
# the alert line
round(mu + z * sigma)  # → 1212
# alert?
10_000 > mu + z * sigma  # → True

In code: TaskBudget.charge counts each step's tokens and raises BudgetExceeded at either limit; TenantSpend.record adds a finished task's cost to its tenant's total and returns a cap alert or an anomaly alert (the formula above).

Unit economics: cost per successful task

Everyday picture. A cheap printer that jams on 40% of pages isn't cheap once you count the wasted paper and the time spent clearing jams.

Level 3: the formula and its symbols

$$ \text{cost per success (retry until it works)} = \frac{c}{p} \qquad \text{cost per task (a person fixes failures)} = c + (1 - p)\,h $$

Symbols

Symbol Meaning here Small model Large model
$c$ model cost of one attempt \$0.002 \$0.010
$p$ chance an attempt succeeds 0.60 0.95
$1/p$ expected number of attempts until one succeeds 1.67 1.05
$h$ cost of a person fixing one failure (illustrative: a few minutes of staff time) \$2.00 \$2.00

In words: if you can simply retry, you pay one attempt's cost for every expected attempt; if failures need a person, every task pays its attempt plus the chance of failure times the cost of the fix.

On the worked example: retrying, the small model costs 0.002 / 0.60 = \$0.0033 per success and the large one 0.010 / 0.95 = \$0.0105, so the small model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 = \$0.802 per task and the large one 0.010 + 0.05 × 2.00 = \$0.110: the "cheap" model is 7x more expensive.

Level 3: in Python

In Python:

def per_success(c, p):
    # retry until it works
    return c / p
def per_task(c, p, h=2.00):
    # a person fixes each failure
    return c + (1 - p) * h
round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4)  # → (0.0033, 0.0105)
round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3)  # → (0.802, 0.11)
# small vs. large, with cleanup
round(per_task(0.002, 0.60) / per_task(0.010, 0.95))  # → 7

With automatic retries the small model is cheaper per success (0.3 vs 1.1 cents); if a person fixes each failure the large model wins, 11 cents vs 80

Reading it: the left pair of bars assumes failures can be retried automatically: the small model wins. The right pair assumes a person has to fix each failure: the large model wins by a mile. Which world you're in decides which model is cheaper, and the model's price per token barely matters in the second one.

In code: cost_per_success_with_retries is $c/p$ and cost_per_success_with_cleanup is $c + (1-p)\,h$.

Putting it together: a 5x plan, in order

Order the levers so the ones that can't hurt quality come first:

  1. Prompt caching: reorder the prompt so the stable prefix is reused. No behaviour change.
  2. Trim tokens: compress tool results and send fewer chunks.
  3. Concise output: ask for shorter answers where length adds nothing.
  4. Route easy tasks to the small model: the first lever that can change quality, so it's validated with evals before and after.
  5. Batch the non-interactive share.

Workload cost falls from 67 to 8 cents as levers stack: caching 2.2x, trimming 3.6x, concise output 3.9x, routing 7.8x, batching 8.4x

Reading it: each bar is the cost of the same 10-task workload after applying every lever up to and including that one; the label is the cumulative reduction. Caching alone roughly halves it, trimming and concise output take it past 3x, routing takes it past 7x, and batching adds the last few percent. The three levers after the baseline (caching, trimming, concise output) are the free wins that don't touch quality.

In code: five_x_plan switches the levers on one at a time, in this order, and reports the cost and cumulative reduction after each.

In 20 seconds

  • Measure cost per successful task, not per call; include retries and human cleanup.
  • Free wins first: prompt caching (stable prefix first), trimming tool output and chunks, concise output, batch APIs for non-interactive work.
  • Then route easy tasks to small models, validated with evals.
  • Run independent tool calls in parallel; stream to improve perceived speed.
  • Budgets per task, spend caps per tenant, and alerts on sudden jumps in tokens per task.
  • Semantic caches can serve confidently wrong answers; use them only on narrow traffic, with guards and a measured wrong-hit rate.

Self-test questions

How would you cut the cost per task by 5x without hurting quality? First measure: cost per successful task, broken down by step, model, input vs. output, and cached vs. uncached tokens, with an eval set to hold quality fixed. Then, in order: restructure prompts for prompt caching; compress tool results and send only the top reranked chunks; ask for concise output; route simple steps (classification, extraction, formatting) to a small model, checking the eval before and after; move anything non-interactive to a batch API; cap steps and tokens per task. On the sample workload here that's about 8x, and every step after the first three is checked against the eval set.

Why can a cheaper model cost more? Because failures cost money too: retries, human cleanup, lost customers. At 60% success with a \$2 human fix, a \$0.002 call costs \$0.80 per task; a \$0.010 call at 95% costs \$0.11.

What are the risks of a semantic cache? Returning a confident, wrong answer to a question that's similar but not the same ("sick days" vs "vacation days"), and serving stale answers after facts change. Mitigate with high thresholds, exact-match guards on IDs, numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit rate on labelled pairs before enabling it.

Your average tokens per task doubled overnight. What do you check? Whether a deploy changed a prompt or a tool (a bigger tool output, a new retrieval setting), whether an agent is looping (the same tool called with the same arguments), and whether the prompt cache hit rate dropped (something volatile moved to the top). Traces make this a lookup rather than a guess (primer.agents.observability).

The papers behind this lesson

  • Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015), https://arxiv.org/abs/1503.02531. Showed how to train a small model to imitate a large one, the technique behind making the cheap model good enough to take more of the routed traffic. annotated companion
  • Chen, Zaharia & Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (2023), https://arxiv.org/abs/2305.05176. Studied prompt adaptation, caching and cascades that try cheap models first and escalate to expensive ones only when needed.

Further reading

on GitHub
   1r"""
   2# Cost and latency: paying less per successful task without getting worse
   3
   4Run: `python -m primer.agents.cost`
   5
   6This lesson builds on tokens and pricing from `primer.ml.inference`, on the
   7context window from `primer.agents.context` and on the release gate from
   8`primer.agents.evals`.
   9
  10## Level 1: The practitioner's guide
  11
  12**In one sentence.** Cost engineering for a model-backed system means paying
  13less per *successful* task, not per call, by taking the levers that can't
  14hurt quality first and checking the ones that can against an eval.
  15
  16**When you need it.** When the bill grows faster than the value, when a
  17latency budget is missed, or before the first big customer arrives and the
  18price per task stops being a rounding error. The tell: you know your spend
  19per month but not your cost per successful task, or you know cost per call
  20but have never counted retries and human clean-up. The lesson's worked
  21example shows why that number, and only that number, decides things: a
  22small model at \$0.002 a call with a 60% success rate beats a large one at
  23\$0.010 and 95% if failures can simply be retried (\$0.0033 against
  24\$0.0105 per success), and loses by 7x if a person has to fix each failure
  25(\$0.802 against \$0.110 per task, at an illustrative \$2 per fix). You don't
  26need any of this for a prototype with ten users a day, and you shouldn't
  27touch the levers that change behaviour (routing, semantic caching) until
  28you have an eval to watch them with. All prices in this lesson are
  29illustrative constants chosen for round arithmetic (a large model at \$5
  30per million input tokens and \$25 per million output, a small one at \$1
  31and \$5); check a provider's current price list for real numbers.
  32
  33**Your options.** Eight levers, from the ones that can't hurt quality to the
  34ones that can:
  35
  36| Option | What it does | What it guarantees | What it costs | Where it lives |
  37|---|---|---|---|---|
  38| Prompt caching | Puts the stable part of the prompt first so the provider reuses its work | No behaviour change; the worked 10,000-token call falls from \$0.0625 to \$0.0265 with 8,000 tokens cached | A small write premium the first time; reordering the prompt | The provider's API |
  39| Trim tokens | Compress tool results, send fewer retrieved chunks, ask for concise output | No change if you only cut what the next step never read; output is the expensive direction | Engineering, and care not to cut what was needed | Your prompt and code |
  40| Parallel tool calls | Run independent tool calls at the same time | Wall-clock time falls to the slowest call (300 ms to 100 ms for three lookups); tokens unchanged | Concurrency in the agent loop | Your agent loop |
  41| Exact response cache | The same question, ignoring case and spaces, returns the stored answer | Never wrong until the facts change | A time-to-live to manage | Your code |
  42| Batch API | Non-interactive work submitted in bulk at about half price | Half price, results within hours | Waiting; only for work nobody is waiting on | The provider's API |
  43| Budgets and alerts | Caps steps and tokens per task and spend per tenant; alerts on a sudden jump | A confused agent stops instead of looping all night | Choosing the limits; a false alarm now and then | Your code |
  44| Routing | Sends easy tasks (classify, extract, format) to a small model and hard ones to a large one | 53% saved on this lesson's ten-task workload, if the router is right | The first lever that can lose quality; needs an eval before and after | Your code, in front of every request |
  45| Semantic cache | Answers a *similar* question from memory, using embeddings to judge similarity | The cheapest possible hit | Confidently wrong answers: on this lesson's pairs, 40% of different-intent questions are served from cache at a 0.95 threshold | Your code, plus an embedding model |
  46
  47A ninth, for later: train or distil a small model on the large model's
  48answers so it can take more of the routed traffic (Hinton et al., 2015).
  49
  50**How to choose.** Measure first, then go down the table in order.
  51
  52- No measurement yet: instrument cost per task by step, by model, input
  53  against output and cached against uncached, and put an eval set in place
  54  to hold quality fixed.
  55- A long, stable system prompt or tool list: prompt caching, today. It is
  56  the only lever that halves a bill without changing a single answer.
  57- Big tool results or many retrieved chunks: trim. Compress to the fields
  58  the next step needs and send the top few reranked chunks.
  59- Several independent lookups per turn: run them concurrently, and stream
  60  the answer so users see progress.
  61- Anything nobody is waiting on (nightly evals, backfills, bulk
  62  classification): batch it.
  63- Most traffic is simple: route, with the eval watching. A router that sends
  64  hard tasks to the small model saves money and quietly loses quality.
  65- Narrow, curated, FAQ-style traffic and nothing else: a semantic cache
  66  with guards (identifiers, numbers and negation must match), a
  67  time-to-live and a measured wrong-hit rate. On everyday questions no
  68  threshold makes it safe.
  69- Whatever you pick, judge it by cost per successful task, with retries and
  70  human clean-up counted in.
  71
  72**What it costs.** The levers stack. On this lesson's ten-task workload the
  73baseline is \$0.671; caching takes it to \$0.311 (2.2x), trimming tool output
  74to \$0.186 (3.6x), concise output to \$0.171 (3.9x), routing to \$0.086 (7.8x)
  75and batching the non-interactive share to \$0.080 (8.4x). The first three
  76are free wins that don't touch quality; routing is the big one and the
  77first that can. Caching is not free on the first request: Anthropic's
  78prompt caching documentation, for one, prices a cache write at 1.25 times
  79the base input price for a five-minute cache (2 times for an hour), reads at
  80a tenth of it or less, and only caches prompts above a per-model minimum of
  81a few hundred to a few thousand tokens. Batching costs time: the same
  82provider quotes a 50% discount with most batches finishing within an hour.
  83Parallel calls and streaming cost no tokens at all; they buy time and
  84perceived speed. Budgets cost the occasional false alarm: this lesson's
  85anomaly rule flags a task using more than the mean plus three standard
  86deviations of recent usage, so a 10,000-token task against a history around
  871,000 trips it at once.
  88
  89**What breaks.**
  90
  91- **Cost per call replacing cost per success.** The cheap model looks
  92  cheaper until failures are priced. Count retries and fixes.
  93- **Routing that loses quality quietly.** Nothing errors; the answers just
  94  get worse. Gate the router with the eval, and re-check when the small
  95  model changes.
  96- **The semantic cache that answers the wrong question.** "How many sick
  97  days do I get?" scores 0.97 against the vacation question and gets the
  98  vacation answer. Near-misses score higher than real paraphrases, so no
  99  threshold separates them.
 100- **Stale cache hits.** The policy changed; the cache didn't. Give every
 101  entry a time-to-live.
 102- **A cache that never hits.** Something volatile (a timestamp, the user's
 103  name) sits at the top of the prompt, so the prefix differs every call.
 104  Stable content first.
 105- **Trimming what was needed.** A tool result cut to 500 tokens loses the
 106  field the next step reads. Compress by field, not by length alone.
 107- **The runaway loop.** An agent calls the same tool with the same
 108  arguments all night. Cap steps and tokens per task; alert on tokens per
 109  task jumping.
 110- **One tenant's surprise invoice.** A shared platform with no per-customer
 111  cap turns one customer's bug into everyone's bill.
 112
 113**In the wild.** Prompt caching and batch processing are standard features
 114of hosted APIs; Anthropic's documentation for both is linked in Further
 115reading and gives the multipliers quoted above. FrugalGPT (Chen, Zaharia and
 116Zou, 2023) named the three families, prompt adaptation, model approximation
 117and cascades that try a cheap model first and escalate, and reported
 118matching the best single model with up to 98% less cost on its benchmarks.
 119Distillation (Hinton, Vinyals and Dean, 2015) is how the small model in the
 120routing table gets good enough to take more traffic. Tooling: LiteLLM is an
 121open-source gateway that puts one interface in front of many providers and
 122adds routing with fallbacks, per-key and per-team budgets, spend tracking
 123and caching; GPTCache is an open-source semantic cache built on embeddings
 124and a vector store with a pluggable similarity evaluator, the design this
 125lesson's `SemanticCache` reproduces in miniature, guards and all. Parallel
 126tool use is part of the tool-calling protocol of the major APIs: the model
 127asks for several tools in one turn and you return all the results in one
 128message.
 129
 130**Go deeper.** Level 2 prices a request symbol by symbol, builds the router,
 131both caches and the parallel loop in plain Python, derives the anomaly
 132alert from a mean and a standard deviation, works the unit economics with
 133every number, and stacks the levers to draw the 8.4x figure. If you only
 134needed to know which lever to pull first, you are done.
 135
 136## Level 2: How it works, from scratch
 137
 138A restaurant doesn't cut costs by buying worse ingredients for every dish.
 139It sends simple orders to the line cook and complex ones to the head chef,
 140pre-chops what every dish shares, doesn't plate food nobody eats, cooks
 141things in parallel, does bulk prep overnight when it's cheaper, and watches
 142the one kitchen station whose bills suddenly spike.
 143
 144Every lever in this lesson is one of those moves. The measure that matters
 145is **cost per successful task**, not cost per call: a cheap attempt that
 146fails still has to be paid for, and so does fixing it.
 147
 148**All prices in this module are illustrative constants** chosen for round
 149arithmetic (a "large" model at \$5 per million input tokens and \$25 per
 150million output tokens; a "small" one at \$1 and \$5). They show the *shape*
 151of the trade-offs; check your provider's current price list for real
 152numbers.
 153
 154## How a request is priced
 155
 156**Everyday picture.** A taxi that charges one rate for the distance to your
 157pickup and a higher rate for the ride itself. Models charge per **token**
 158(a word piece, about 4 characters of English): one rate for tokens you send
 159(**input**) and a higher rate for tokens the model writes (**output**),
 160because writing happens one token at a time (see `primer.ml.inference`).
 161
 162**Worked example.** A large-model call with 10,000 input tokens and 500
 163output tokens: 10,000 × \$5 / 1,000,000 = \$0.05 for input, plus
 164500 × \$25 / 1,000,000 = \$0.0125 for output, total **\$0.0625**. If 8,000 of
 165those input tokens are a stable prefix served from the **prompt cache**
 166(reused work from an earlier identical beginning, billed here at 10% of the
 167input price), input becomes 8,000 × \$0.50/M + 2,000 × \$5/M = \$0.014, and
 168the call costs **\$0.0265**, 58% less.
 169
 170$$
 171\text{cost} = \frac{c \cdot \rho\, p_{\text{in}} + (n_{\text{in}} - c)\, p_{\text{in}} + n_{\text{out}}\, p_{\text{out}}}{10^6}
 172$$
 173
 174**Symbols**
 175
 176| Symbol | Meaning here | Worked example |
 177|---|---|---|
 178| $n_{\text{in}}$ | input tokens sent | 10,000 |
 179| $c$ | input tokens served from the prompt cache | 8,000 |
 180| $n_{\text{out}}$ | output tokens generated | 500 |
 181| $p_{\text{in}}, p_{\text{out}}$ | price per million input / output tokens | \$5, \$25 |
 182| $\rho$ | cached-read price as a fraction of normal input price (Greek *rho*) | 0.1 |
 183| $10^6$ | prices are per million tokens | |
 184
 185**In words:** cached input at its discounted rate, plus the rest of the
 186input at full rate, plus output at the output rate, all per million.
 187
 188**On the worked example:** (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ =
 189(4,000 + 10,000 + 12,500) / 10⁶ = \$0.0265.
 190
 191**In Python:**
 192
 193```python
 194n_in, c, n_out = 10_000, 8_000, 500
 195p_in, p_out, rho = 5, 25, 0.1
 196cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6
 197round(cost, 4)  # → 0.0265
 198```
 199
 200Two facts fall out: output tokens cost several times more than input (ask
 201for concise answers), and a cached prefix is nearly free (put stable content
 202first; see `primer.agents.context`). Caches usually charge a small premium
 203the first time a prefix is written; this module ignores it for simplicity.
 204
 205**In code:** `request_cost` is the formula above, reading each model's
 206`Price` from `PRICES`.
 207
 208## Route each task to the cheapest model that can do it
 209
 210**Everyday picture.** A hospital triage nurse: sprained ankles go to the
 211nurse practitioner, chest pains to the cardiologist. Nobody sends every
 212patient to the most expensive specialist.
 213
 214**Worked example.** The sample workload is 10 tasks: 7 simple (classify a
 215ticket, extract a date) and 3 complex (plan a migration). All on the large
 216model it costs \$0.671. Routing the 7 simple ones to the small model, which
 217is 5x cheaper per token, brings it to \$0.314: **53% saved**, with the hard
 218tasks still on the strong model.
 219
 220```mermaid
 221flowchart LR
 222  T[Incoming task] --> R{Router<br/>rules or a small classifier}
 223  R -->|classify, extract,<br/>format, short| S[Small, fast model]
 224  R -->|plan, analyze,<br/>multi-step reasoning| L[Large model]
 225  S --> O[Result]
 226  L --> O
 227```
 228
 229**Reading it:** the router is cheap (a few rules here; often a small
 230classifier) and sits in front of every request. The whole saving comes from
 231the left branch: most traffic in real systems is simple, and simple work
 232runs well on small models. Validate the router with evals
 233(`primer.agents.evals`): a router that sends hard tasks to the small model
 234saves money and quietly loses quality.
 235
 236**In code:** `route` is the rule-based router; `workload_cost` prices a list
 237of `WorkItem` tasks with any levers switched on; `routing_savings` compares
 238all-large against routed on `SAMPLE_WORKLOAD`.
 239
 240## Response caching and semantic caching
 241
 242**Everyday picture.** A receptionist who's been asked "what's the wifi
 243password?" a hundred times just answers from memory. That's an **exact
 244response cache**: the same question (ignoring case and spaces) returns the
 245stored answer with no model call. A **semantic cache** goes further and
 246answers *similar* questions from memory, using embeddings (vectors where
 247closeness means similar meaning, see `primer.ml.embeddings`) to decide
 248"similar". That's where the receptionist starts giving the vacation-policy
 249answer to someone asking about sick leave.
 250
 251**Worked example.** With the toy embedder in `primer.common.embedder`:
 252
 253| Stored question | New question | Similarity | Same intent? |
 254|---|---|---|---|
 255| How do I reset my password? | I forgot my password, how do I recover it? | 0.88 | yes |
 256| How many vacation days do I get? | How many sick days do I get? | **0.97** | **no** |
 257| Request a new laptop | Request a new monitor | **0.96** | **no** |
 258| What does ERR-4012 mean? | What does ERR-4013 mean? | 0.56 | no |
 259
 260The two near-misses score *higher* than a genuine paraphrase. No single
 261threshold separates them. Two cheap **guards** help: identifiers and
 262numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match
 263("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which
 264share a topic and differ in a single word.
 265
 266```mermaid
 267flowchart TD
 268  Q[New question] --> E{Exact match<br/>in response cache?}
 269  E -->|yes| A1[Return stored answer]
 270  E -->|no| V[Embed and find the<br/>nearest stored question]
 271  V --> T{Similarity above<br/>threshold?}
 272  T -->|no| M[Call the model]
 273  T -->|yes| G{IDs, numbers and<br/>negation identical?<br/>entry not expired?}
 274  G -->|no| M
 275  G -->|yes| A2[Return stored answer]
 276  M --> S[Store the new answer]
 277```
 278
 279**Reading it:** the exact cache is checked first because it's free and never
 280wrong. The semantic path adds two gates: the similarity threshold, and the
 281guards. An expired entry (older than its **TTL**, time to live) is treated
 282as a miss, so answers about changing facts don't go stale forever.
 283
 284![No threshold separates them: wrong hits stay at 40 to 60% through 0.95, often above the paraphrase hit rate, and reach 0 only where paraphrase hits do too](figures/primer.agents.cost.semantic_cache.svg)
 285
 286**Reading it:** the x-axis is the similarity threshold; the blue line is the
 287share of genuine paraphrases answered from cache (good), and the red line is
 288the share of different-intent questions answered from cache (a wrong answer
 289served confidently). Lowering the threshold raises both, but look where the
 290lines sit: from 0.4 upward the red line is above the blue one, so the cache
 291serves more wrong answers than right ones (60% of different-intent questions
 292against at most 57% of paraphrases). Even at 0.95 it still answers 40% of
 293them wrongly and only 14% of paraphrases, because "sick" vs "vacation"
 294scores 0.97. On everyday questions no threshold makes semantic caching safe;
 295it's safe only for narrow, curated FAQ-style traffic, with guards, a TTL,
 296and a measured wrong-hit rate.
 297
 298**In code:** `ResponseCache` is the exact cache. `SemanticCache` is the
 299semantic path: `SemanticCache.lookup` skips expired entries and entries whose
 300identifiers or negation differ, then returns the nearest survivor, and
 301`SemanticCache.get` applies the threshold. `semantic_cache_sweep` draws the
 302figure from the labelled pairs in `CACHE_PAIRS`.
 303
 304## Trim tokens
 305
 306**Everyday picture.** Don't photocopy the whole binder when the colleague
 307needs one page. Compress tool results to the fields the next step needs
 308(`primer.agents.context.compress_tool_output`), send the top few reranked
 309chunks instead of dozens, remove repeated boilerplate from system prompts,
 310and ask for concise output, because output is the expensive direction.
 311
 312**In code:** `workload_cost` models trimming as cutting each task's tool
 313output to 500 tokens, and concise output as cutting long answers by a third.
 314
 315## Run independent tool calls at the same time
 316
 317**Everyday picture.** Boil the pasta while the sauce simmers. Cooking them
 318one after the other takes the sum of the times; together, the time of the
 319slowest.
 320
 321**Worked example.** Three independent lookups of 100 ms each: sequentially
 322about 300 ms; concurrently about 100 ms.
 323
 324```mermaid
 325sequenceDiagram
 326  participant A as Agent
 327  participant T1 as Weather API
 328  participant T2 as Calendar API
 329  participant T3 as CRM API
 330  Note over A,T3: Sequential: about 300 ms
 331  A->>T1: call
 332  T1-->>A: result
 333  A->>T2: call
 334  T2-->>A: result
 335  A->>T3: call
 336  T3-->>A: result
 337  Note over A,T3: Parallel: about 100 ms
 338  par
 339    A->>T1: call
 340  and
 341    A->>T2: call
 342  and
 343    A->>T3: call
 344  end
 345  T1-->>A: result
 346  T2-->>A: result
 347  T3-->>A: result
 348```
 349
 350**Reading it:** in the top half each call waits for the previous result; in
 351the bottom half the three requests leave together and the agent waits only
 352for the slowest. Models can ask for several tools in one turn (parallel tool
 353use); run them concurrently and return all results in one message. Parallel
 354calls cut wall-clock time, not tokens.
 355
 356![Three 100 ms tool calls take 300 ms end to end when run one after another, but only 100 ms when started together](figures/primer.agents.cost.parallel.svg)
 357
 358**Reading it:** each bar is one 100 ms tool call on a shared time axis. The
 359sequential calls stack end to end, 300 ms in all; the parallel ones start
 360together, so the whole batch finishes in the time of one. Run the lesson to
 361measure it with `asyncio`: the real timings land within a few milliseconds
 362of these bars.
 363
 364**In code:** `run_sequential` awaits each call before starting the next;
 365`run_parallel` starts them all with Python's asyncio gather and waits once.
 366
 367**Streaming** (showing tokens as they're generated) doesn't reduce total
 368time either, but users see progress immediately, which changes how fast the
 369system *feels*.
 370
 371## Batch what isn't interactive
 372
 373**Everyday picture.** Sending the laundry out to be done overnight at half
 374price instead of waiting at the express counter. Providers offer **batch
 375APIs**: submit many requests, get results within hours (often within 24),
 376typically at about half the price. Use them for anything no one is waiting
 377on: nightly evals, backfilling document processing, bulk classification.
 378
 379**In code:** `batch_cost` prices a list of requests at the batch discount.
 380
 381## Budgets and alerts
 382
 383**Everyday picture.** A prepaid card with a hard limit, and a bank that
 384texts you when a purchase looks nothing like your usual spending.
 385
 386A **task budget** caps steps and tokens per task, so a confused agent stops
 387instead of looping all night. A **tenant spend cap** (a tenant is one
 388customer organisation on a shared platform) stops one customer's runaway
 389usage from becoming a surprise invoice. An **anomaly alert** fires when a
 390task uses far more tokens than usual:
 391
 392$$
 393\text{alert if } x > \mu + z\,\sigma
 394$$
 395
 396**Symbols**
 397
 398| Symbol | Meaning here | Worked example |
 399|---|---|---|
 400| $x$ | tokens used by the task just finished | 10,000 |
 401| $\mu$ | mean tokens per task over this tenant's recent history (Greek *mu*) | 1,000 |
 402| $\sigma$ | standard deviation of that history: the typical distance from the mean (Greek *sigma*) | ≈ 71 |
 403| $z$ | how many standard deviations count as unusual | 3 |
 404
 405**In words:** alert when a task uses more tokens than the usual amount plus
 406three times the usual spread.
 407
 408**On the worked example:** history 1000, 1100, 900, 1050, 950, 1000 has mean
 4091,000 and standard deviation ≈ 71 (the squared distances from the mean add
 410up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7),
 411so the line is 1,000 + 3 × 70.7 ≈ 1,212.
 412A 10,000-token task is far above it: alert. Such jumps usually mean a loop
 413or a bad deploy.
 414
 415**In Python:**
 416
 417```python
 418import statistics
 419history = [1000, 1100, 900, 1050, 950, 1000]
 420mu = statistics.mean(history)
 421# divides by 6 - 1, as above
 422sigma = statistics.stdev(history)
 423mu, round(sigma, 1)  # → (1000, 70.7)
 424z = 3
 425# the alert line
 426round(mu + z * sigma)  # → 1212
 427# alert?
 42810_000 > mu + z * sigma  # → True
 429```
 430
 431**In code:** `TaskBudget.charge` counts each step's tokens and raises
 432`BudgetExceeded` at either limit; `TenantSpend.record` adds a finished task's
 433cost to its tenant's total and returns a cap alert or an anomaly alert (the
 434formula above).
 435
 436## Unit economics: cost per successful task
 437
 438**Everyday picture.** A cheap printer that jams on 40% of pages isn't
 439cheap once you count the wasted paper and the time spent clearing jams.
 440
 441$$
 442\text{cost per success (retry until it works)} = \frac{c}{p}
 443\qquad
 444\text{cost per task (a person fixes failures)} = c + (1 - p)\,h
 445$$
 446
 447**Symbols**
 448
 449| Symbol | Meaning here | Small model | Large model |
 450|---|---|---|---|
 451| $c$ | model cost of one attempt | \$0.002 | \$0.010 |
 452| $p$ | chance an attempt succeeds | 0.60 | 0.95 |
 453| $1/p$ | expected number of attempts until one succeeds | 1.67 | 1.05 |
 454| $h$ | cost of a person fixing one failure (illustrative: a few minutes of staff time) | \$2.00 | \$2.00 |
 455
 456**In words:** if you can simply retry, you pay one attempt's cost for every
 457expected attempt; if failures need a person, every task pays its attempt
 458plus the chance of failure times the cost of the fix.
 459
 460**On the worked example:** retrying, the small model costs 0.002 / 0.60 =
 461\$0.0033 per success and the large one 0.010 / 0.95 = \$0.0105, so the small
 462model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 =
 463**\$0.802** per task and the large one 0.010 + 0.05 × 2.00 = **\$0.110**: the
 464"cheap" model is 7x more expensive.
 465
 466**In Python:**
 467
 468```python
 469def per_success(c, p):
 470    # retry until it works
 471    return c / p
 472def per_task(c, p, h=2.00):
 473    # a person fixes each failure
 474    return c + (1 - p) * h
 475round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4)  # → (0.0033, 0.0105)
 476round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3)  # → (0.802, 0.11)
 477# small vs. large, with cleanup
 478round(per_task(0.002, 0.60) / per_task(0.010, 0.95))  # → 7
 479```
 480
 481![With automatic retries the small model is cheaper per success (0.3 vs 1.1 cents); if a person fixes each failure the large model wins, 11 cents vs 80](figures/primer.agents.cost.unit_economics.svg)
 482
 483**Reading it:** the left pair of bars assumes failures can be retried
 484automatically: the small model wins. The right pair assumes a person has to
 485fix each failure: the large model wins by a mile. Which world you're in
 486decides which model is cheaper, and the model's price per token barely
 487matters in the second one.
 488
 489**In code:** `cost_per_success_with_retries` is $c/p$ and
 490`cost_per_success_with_cleanup` is $c + (1-p)\,h$.
 491
 492## Putting it together: a 5x plan, in order
 493
 494Order the levers so the ones that can't hurt quality come first:
 495
 4961. **Prompt caching**: reorder the prompt so the stable prefix is reused.
 497   No behaviour change.
 4982. **Trim tokens**: compress tool results and send fewer chunks.
 4993. **Concise output**: ask for shorter answers where length adds nothing.
 5004. **Route easy tasks to the small model**: the first lever that *can*
 501   change quality, so it's validated with evals before and after.
 5025. **Batch** the non-interactive share.
 503
 504![Workload cost falls from 67 to 8 cents as levers stack: caching 2.2x, trimming 3.6x, concise output 3.9x, routing 7.8x, batching 8.4x](figures/primer.agents.cost.five_x.svg)
 505
 506**Reading it:** each bar is the cost of the same 10-task workload after
 507applying every lever up to and including that one; the label is the
 508cumulative reduction. Caching alone roughly halves it, trimming and concise
 509output take it past 3x, routing takes it past 7x, and batching adds the
 510last few percent. The three levers after the baseline (caching, trimming,
 511concise output) are the free wins that don't touch quality.
 512
 513**In code:** `five_x_plan` switches the levers on one at a time, in this
 514order, and reports the cost and cumulative reduction after each.
 515
 516## In 20 seconds
 517- Measure cost per *successful* task, not per call; include retries and
 518  human cleanup.
 519- Free wins first: prompt caching (stable prefix first), trimming tool
 520  output and chunks, concise output, batch APIs for non-interactive work.
 521- Then route easy tasks to small models, validated with evals.
 522- Run independent tool calls in parallel; stream to improve perceived
 523  speed.
 524- Budgets per task, spend caps per tenant, and alerts on sudden jumps in
 525  tokens per task.
 526- Semantic caches can serve confidently wrong answers; use them only on
 527  narrow traffic, with guards and a measured wrong-hit rate.
 528
 529## Self-test questions
 530
 531**How would you cut the cost per task by 5x without hurting quality?**
 532First measure: cost per successful task, broken down by step, model, input
 533vs. output, and cached vs. uncached tokens, with an eval set to hold
 534quality fixed. Then, in order: restructure prompts for prompt caching;
 535compress tool results and send only the top reranked chunks; ask for
 536concise output; route simple steps (classification, extraction,
 537formatting) to a small model, checking the eval before and after; move
 538anything non-interactive to a batch API; cap steps and tokens per task.
 539On the sample workload here that's about 8x, and every step after the
 540first three is checked against the eval set.
 541
 542**Why can a cheaper model cost more?**
 543Because failures cost money too: retries, human cleanup, lost customers.
 544At 60% success with a \$2 human fix, a \$0.002 call costs \$0.80 per task;
 545a \$0.010 call at 95% costs \$0.11.
 546
 547**What are the risks of a semantic cache?**
 548Returning a confident, wrong answer to a question that's similar but not
 549the same ("sick days" vs "vacation days"), and serving stale answers after
 550facts change. Mitigate with high thresholds, exact-match guards on IDs,
 551numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit
 552rate on labelled pairs before enabling it.
 553
 554**Your average tokens per task doubled overnight. What do you check?**
 555Whether a deploy changed a prompt or a tool (a bigger tool output, a new
 556retrieval setting), whether an agent is looping (the same tool called with
 557the same arguments), and whether the prompt cache hit rate dropped
 558(something volatile moved to the top). Traces make this a lookup rather
 559than a guess (`primer.agents.observability`).
 560
 561## The papers behind this lesson
 562
 563- Hinton, Vinyals & Dean, *Distilling the Knowledge in a Neural Network*
 564  (2015), https://arxiv.org/abs/1503.02531. Showed how to train a small
 565  model to imitate a large one, the technique behind making the cheap model
 566  good enough to take more of the routed traffic.
 567  [annotated companion](../../papers/distillation.html)
 568- Chen, Zaharia & Zou, *FrugalGPT: How to Use Large Language Models While
 569  Reducing Cost and Improving Performance* (2023),
 570  https://arxiv.org/abs/2305.05176. Studied prompt adaptation, caching and
 571  cascades that try cheap models first and escalate to expensive ones only
 572  when needed.
 573
 574## Further reading
 575- Anthropic prompt caching docs: https://docs.claude.com/en/docs/build-with-claude/prompt-caching
 576- Anthropic Message Batches docs: https://docs.claude.com/en/docs/build-with-claude/batch-processing
 577- Anthropic, parallel tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
 578- Python `asyncio.gather`: https://docs.python.org/3/library/asyncio-task.html#asyncio.gather
 579"""
 580
 581from __future__ import annotations
 582
 583import asyncio
 584import re
 585import statistics
 586import time
 587from dataclasses import dataclass, field
 588from typing import Any, Callable
 589
 590import numpy as np
 591
 592from primer._show import banner, say, table, takeaway
 593from primer.common.embedder import ConceptEmbedder
 594
 595# ---------------------------------------------------------------------------
 596# 1. Pricing (ILLUSTRATIVE constants, not any vendor's price list)
 597# ---------------------------------------------------------------------------
 598
 599
 600@dataclass(frozen=True)
 601class Price:
 602    input_per_m: float  # dollars per million input tokens
 603    output_per_m: float  # dollars per million output tokens
 604
 605
 606PRICES: dict[str, Price] = {"small": Price(1.0, 5.0), "large": Price(5.0, 25.0)}
 607CACHE_READ_MULTIPLIER = 0.1  # cached input billed at 10% of the input price
 608BATCH_MULTIPLIER = 0.5  # batch interface at half price
 609
 610
 611def request_cost(model: str, input_tokens: int, output_tokens: int, cached_tokens: int = 0) -> float:
 612    """Dollar cost of one call. `cached_tokens` of the input are billed at the cached rate."""
 613    p = PRICES[model]
 614    uncached = input_tokens - cached_tokens
 615    return (cached_tokens * CACHE_READ_MULTIPLIER * p.input_per_m + uncached * p.input_per_m + output_tokens * p.output_per_m) / 1e6
 616
 617
 618def batch_cost(requests: list[tuple[str, int, int]]) -> float:
 619    """Cost of (model, input, output) requests submitted through a batch interface."""
 620    return BATCH_MULTIPLIER * sum(request_cost(m, i, o) for m, i, o in requests)
 621
 622
 623# ---------------------------------------------------------------------------
 624# 2. Routing
 625# ---------------------------------------------------------------------------
 626
 627# Words that suggest multi-step reasoning. A real router is often a small
 628# classifier trained on labelled traffic; rules are the transparent start.
 629HARD_SIGNALS = ("plan", "analy", "design", "strategy", "compare", "debug", "reason", "multi-step", "trade-off", "why")
 630
 631
 632def route(task_text: str, max_simple_chars: int = 400) -> str:
 633    """Return "small" or "large" for a task."""
 634    text = task_text.lower()
 635    if len(text) > max_simple_chars or any(s in text for s in HARD_SIGNALS):
 636        return "large"
 637    return "small"
 638
 639
 640@dataclass
 641class WorkItem:
 642    text: str
 643    output_tokens: int
 644    interactive: bool = True
 645    # Input is split so the levers have something to act on.
 646    stable_prefix: int = 8_000  # system prompt + tool definitions, identical every call
 647    tool_output: int = 3_000  # verbose tool results that could be compressed
 648    rest: int = 1_000  # the question and its specific context
 649
 650
 651SAMPLE_WORKLOAD: list[WorkItem] = (
 652    [WorkItem("Classify this ticket as billing, technical or other", 150, interactive=False) for _ in range(4)]
 653    + [WorkItem("Extract the invoice date and total from this email", 150) for _ in range(3)]
 654    + [WorkItem("Plan the migration of our billing system and analyze the risks of each step", 600) for _ in range(3)]
 655)
 656
 657
 658def workload_cost(work: list[WorkItem], cache: bool = False, trim: bool = False, concise: bool = False,
 659                  routed: bool = False, batch: bool = False) -> float:
 660    """Cost of a workload with any combination of levers switched on."""
 661    total = 0.0
 662    for w in work:
 663        tool = 500 if trim else w.tool_output  # compress tool output to the needed fields
 664        out = int(w.output_tokens * 2 / 3) if concise and w.output_tokens > 300 else w.output_tokens
 665        n_in = w.stable_prefix + tool + w.rest
 666        model = route(w.text) if routed else "large"
 667        c = request_cost(model, n_in, out, cached_tokens=w.stable_prefix if cache else 0)
 668        if batch and not w.interactive:
 669            c *= BATCH_MULTIPLIER
 670        total += c
 671    return total
 672
 673
 674def routing_savings(work: list[WorkItem] | None = None) -> dict[str, float]:
 675    work = work or SAMPLE_WORKLOAD
 676    all_large = workload_cost(work)
 677    routed = workload_cost(work, routed=True)
 678    return {"all_large": all_large, "routed": routed, "savings": 1 - routed / all_large}
 679
 680
 681# ---------------------------------------------------------------------------
 682# 3. Response cache and semantic cache
 683# ---------------------------------------------------------------------------
 684
 685
 686def _normalize(q: str) -> str:
 687    return " ".join(q.lower().split())
 688
 689
 690class ResponseCache:
 691    """Exact-match cache: free and never wrong, but misses every paraphrase."""
 692
 693    def __init__(self) -> None:
 694        self._store: dict[str, str] = {}
 695
 696    def put(self, question: str, answer: str) -> None:
 697        self._store[_normalize(question)] = answer
 698
 699    def get(self, question: str) -> str | None:
 700        return self._store.get(_normalize(question))
 701
 702
 703_NEGATION = re.compile(r"\b(not|no|never|don't|dont|doesn't|isn't|can't|won't|cannot)\b")
 704
 705
 706def _identifiers(text: str) -> set[str]:
 707    """Tokens containing a digit: order ids, error codes, amounts, dates."""
 708    return {t for t in re.findall(r"[a-z0-9]+(?:-[a-z0-9]+)*", text.lower()) if any(c.isdigit() for c in t)}
 709
 710
 711def _negated(text: str) -> bool:
 712    return bool(_NEGATION.search(text.lower()))
 713
 714
 715@dataclass
 716class _Entry:
 717    vector: np.ndarray
 718    question: str
 719    answer: str
 720    stored_at: float
 721
 722
 723class SemanticCache:
 724    """Answer similar questions from memory, with guards against the obvious wrong hits.
 725
 726    Guards: identifiers and numbers must match exactly, negation must match,
 727    and entries expire after `ttl_seconds`. None of this can separate
 728    "sick days" from "vacation days"; only narrow scoping and a measured
 729    wrong-hit rate can.
 730    """
 731
 732    def __init__(self, embedder: Any = None, threshold: float = 0.9, ttl_seconds: float | None = None,
 733                 clock: Callable[[], float] = time.monotonic):
 734        self.embedder = embedder or ConceptEmbedder()
 735        self.threshold = threshold
 736        self.ttl = ttl_seconds
 737        self.clock = clock
 738        self.entries: list[_Entry] = []
 739
 740    def put(self, question: str, answer: str) -> None:
 741        self.entries.append(_Entry(self.embedder.encode(question), question, answer, self.clock()))
 742
 743    def lookup(self, question: str) -> tuple[_Entry | None, float]:
 744        """Nearest stored entry that passes the guards, and its similarity."""
 745        q = self.embedder.encode(question)
 746        best, best_sim = None, -1.0
 747        for e in self.entries:
 748            if self.ttl is not None and self.clock() - e.stored_at > self.ttl:
 749                continue
 750            if _identifiers(e.question) != _identifiers(question) or _negated(e.question) != _negated(question):
 751                continue
 752            sim = float(e.vector @ q)  # vectors are unit length, so dot product = cosine similarity
 753            if sim > best_sim:
 754                best, best_sim = e, sim
 755        return best, best_sim
 756
 757    def get(self, question: str) -> str | None:
 758        entry, sim = self.lookup(question)
 759        return entry.answer if entry is not None and sim >= self.threshold else None
 760
 761
 762# (stored question, new question, same intent?)
 763CACHE_PAIRS: list[tuple[str, str, bool]] = [
 764    ("How do I reset my password?", "I forgot my password, how do I recover it?", True),
 765    ("How do I reset my password?", "How do I reset my VPN password?", False),
 766    ("How many vacation days do I get?", "How much PTO do I get per year?", True),
 767    ("How many vacation days do I get?", "How many sick days do I get?", False),
 768    ("How do I get reimbursed for car mileage?", "automobile expense reimbursement", True),
 769    ("Cancel my order", "Don't cancel my order", False),
 770    ("What does ERR-4012 mean?", "What does ERR-4013 mean?", False),
 771    ("How do I set up the VPN?", "VPN setup steps", True),
 772    ("Report a phishing email", "How do I report a suspicious email?", True),
 773    ("Request a new laptop", "Request a new monitor", False),
 774    ("What is the travel per-diem?", "How much is the daily meal per-diem when traveling?", True),
 775    ("Printer not working", "printing fails on floor printer", True),
 776]
 777
 778
 779def semantic_cache_sweep(thresholds: list[float]) -> list[dict[str, float]]:
 780    """For each threshold: share of paraphrases served from cache, share of near-misses served wrongly."""
 781    out = []
 782    for th in thresholds:
 783        useful = wrong = 0
 784        for stored, new, same in CACHE_PAIRS:
 785            cache = SemanticCache(threshold=th)
 786            cache.put(stored, "cached answer")
 787            hit = cache.get(new) is not None
 788            useful += hit and same
 789            wrong += hit and not same
 790        n_same = sum(1 for p in CACHE_PAIRS if p[2])
 791        out.append({"threshold": th, "useful_hit_rate": useful / n_same, "wrong_hit_rate": wrong / (len(CACHE_PAIRS) - n_same)})
 792    return out
 793
 794
 795# ---------------------------------------------------------------------------
 796# 4. Parallel tool calls
 797# ---------------------------------------------------------------------------
 798
 799
 800async def _fake_tool(i: int, delay: float, t0: float, log: list) -> str:
 801    start = time.perf_counter() - t0
 802    await asyncio.sleep(delay)  # stands in for a network call
 803    log.append((i, start, time.perf_counter() - t0))
 804    return f"tool-{i}"
 805
 806
 807async def run_sequential(delays: list[float]) -> dict[str, Any]:
 808    t0, log = time.perf_counter(), []
 809    results = [await _fake_tool(i, d, t0, log) for i, d in enumerate(delays)]
 810    return {"seconds": time.perf_counter() - t0, "results": results, "spans": sorted(log)}
 811
 812
 813async def run_parallel(delays: list[float]) -> dict[str, Any]:
 814    t0, log = time.perf_counter(), []
 815    # gather starts every call before awaiting any, and returns results in call order.
 816    results = await asyncio.gather(*(_fake_tool(i, d, t0, log) for i, d in enumerate(delays)))
 817    return {"seconds": time.perf_counter() - t0, "results": list(results), "spans": sorted(log)}
 818
 819
 820# ---------------------------------------------------------------------------
 821# 5. Budgets and spend alerts
 822# ---------------------------------------------------------------------------
 823
 824
 825class BudgetExceeded(RuntimeError):
 826    pass
 827
 828
 829@dataclass
 830class TaskBudget:
 831    """Hard per-task limits. The agent loop calls charge() before every model call."""
 832
 833    max_steps: int
 834    max_tokens: int
 835    steps: int = 0
 836    tokens: int = 0
 837
 838    def charge(self, tokens: int) -> None:
 839        if self.steps + 1 > self.max_steps:
 840            raise BudgetExceeded(f"step limit {self.max_steps} reached")
 841        if self.tokens + tokens > self.max_tokens:
 842            raise BudgetExceeded(f"token limit {self.max_tokens} would be exceeded ({self.tokens} + {tokens})")
 843        self.steps += 1
 844        self.tokens += tokens
 845
 846
 847@dataclass
 848class TenantSpend:
 849    """Per-tenant monthly spend cap plus an anomaly alert on tokens per task."""
 850
 851    monthly_cap: float
 852    z: float = 3.0
 853    min_history: int = 5
 854    spent: dict[str, float] = field(default_factory=dict)
 855    history: dict[str, list[int]] = field(default_factory=dict)
 856
 857    def record(self, tenant: str, cost: float, tokens: int) -> dict[str, str]:
 858        """Record one finished task; return alerts keyed by kind ("cap", "anomaly")."""
 859        alerts: dict[str, str] = {}
 860        self.spent[tenant] = self.spent.get(tenant, 0.0) + cost
 861        if self.spent[tenant] > self.monthly_cap:
 862            alerts["cap"] = f"{tenant} spent {self.spent[tenant]:.2f} of a {self.monthly_cap:.2f} cap"
 863        hist = self.history.setdefault(tenant, [])
 864        if len(hist) >= self.min_history:
 865            mu, sigma = statistics.mean(hist), statistics.stdev(hist)
 866            if tokens > mu + self.z * max(sigma, 1.0):
 867                alerts["anomaly"] = f"{tokens} tokens vs usual {mu:.0f} ± {sigma:.0f}"
 868        hist.append(tokens)
 869        return alerts
 870
 871
 872# ---------------------------------------------------------------------------
 873# 6. Unit economics and the plan
 874# ---------------------------------------------------------------------------
 875
 876
 877def cost_per_success_with_retries(cost_per_attempt: float, success_rate: float) -> float:
 878    """Retry until success: expected attempts are 1/p."""
 879    return cost_per_attempt / success_rate
 880
 881
 882def cost_per_success_with_cleanup(cost_per_attempt: float, success_rate: float, human_fix_cost: float) -> float:
 883    """One attempt per task; a person fixes each failure."""
 884    return cost_per_attempt + (1 - success_rate) * human_fix_cost
 885
 886
 887def five_x_plan(work: list[WorkItem] | None = None) -> list[dict[str, Any]]:
 888    """Apply the levers cumulatively, free ones first. Returns cost after each."""
 889    work = work or SAMPLE_WORKLOAD
 890    levers = [
 891        ("baseline: everything on the large model", {}),
 892        ("prompt caching", {"cache": True}),
 893        ("trim tool output", {"cache": True, "trim": True}),
 894        ("concise output", {"cache": True, "trim": True, "concise": True}),
 895        ("route easy tasks to the small model", {"cache": True, "trim": True, "concise": True, "routed": True}),
 896        ("batch non-interactive work", {"cache": True, "trim": True, "concise": True, "routed": True, "batch": True}),
 897    ]
 898    base = workload_cost(work)
 899    steps = []
 900    for name, kw in levers:
 901        c = workload_cost(work, **kw)
 902        steps.append({"lever": name, "cost": c, "factor": base / c})
 903    return steps
 904
 905
 906# ---------------------------------------------------------------------------
 907# Figures and demo
 908# ---------------------------------------------------------------------------
 909
 910
 911def figures() -> dict[str, Any]:
 912    import matplotlib
 913
 914    matplotlib.use("Agg")
 915    import matplotlib.pyplot as plt
 916
 917    figs: dict[str, Any] = {}
 918
 919    # 1. Unit economics.
 920    fig, ax = plt.subplots(figsize=(6.5, 3.5))
 921    labels = ["retry until success", "person fixes failures"]
 922    small = [cost_per_success_with_retries(0.002, 0.60), cost_per_success_with_cleanup(0.002, 0.60, 2.0)]
 923    large = [cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)]
 924    x = np.arange(2)
 925    ax.bar(x - 0.2, small, 0.4, label="small: $0.002/attempt, 60% success", color="#c44e52")
 926    ax.bar(x + 0.2, large, 0.4, label="large: $0.010/attempt, 95% success", color="#4c72b0")
 927    for xi, (s, lg) in enumerate(zip(small, large)):
 928        ax.text(xi - 0.2, s, f"${s:.3f}", ha="center", va="bottom", fontsize=8)
 929        ax.text(xi + 0.2, lg, f"${lg:.3f}", ha="center", va="bottom", fontsize=8)
 930    ax.set_yscale("log")
 931    ax.set_xticks(x, labels)
 932    ax.set_ylabel("dollars per successful task (log scale)")
 933    ax.set_title("The cheap model is only cheap if failures are free")
 934    ax.legend(fontsize=8, loc="upper left")
 935    fig.tight_layout()
 936    figs["unit_economics"] = fig
 937
 938    # 2. The 5x plan.
 939    steps = five_x_plan()
 940    fig, ax = plt.subplots(figsize=(7.5, 3.8))
 941    ax.bar(range(len(steps)), [s["cost"] for s in steps], color=["#8c8c8c"] + ["#4c72b0"] * 3 + ["#dd8452"] * 2)
 942    for i, s in enumerate(steps):
 943        ax.text(i, s["cost"], f"{s['factor']:.1f}x", ha="center", va="bottom", fontsize=9)
 944    ax.set_xticks(range(len(steps)), ["baseline", "+ caching", "+ trim", "+ concise", "+ routing", "+ batch"])
 945    ax.set_ylabel("workload cost ($)")
 946    ax.set_title("Levers applied in order (blue: can't change quality; orange: validate with evals)")
 947    fig.tight_layout()
 948    figs["five_x"] = fig
 949
 950    # 3. Semantic cache threshold sweep.
 951    ths = list(np.round(np.arange(0.3, 1.001, 0.025), 3))
 952    sweep = semantic_cache_sweep(ths)
 953    fig, ax = plt.subplots(figsize=(6.5, 3.5))
 954    ax.plot(ths, [s["useful_hit_rate"] for s in sweep], color="#4c72b0", label="paraphrases served from cache (good)")
 955    ax.plot(ths, [s["wrong_hit_rate"] for s in sweep], color="#c44e52", label="different questions served from cache (wrong)")
 956    ax.set_xlabel("similarity threshold")
 957    ax.set_ylabel("share of pairs")
 958    ax.set_title("Semantic cache: no threshold separates them cleanly")
 959    ax.legend(fontsize=8)
 960    fig.tight_layout()
 961    figs["semantic_cache"] = fig
 962
 963    # 4. Parallel vs sequential timeline: the schedule each approach follows. Drawn from the delays, not
 964    # a stopwatch, so the site doesn't change with a millisecond of jitter; the demo measures the real thing.
 965    delays_ms = [100, 100, 100]
 966    sequential = [(i, sum(delays_ms[:i]), sum(delays_ms[: i + 1])) for i in range(len(delays_ms))]
 967    parallel = [(i, 0, d) for i, d in enumerate(delays_ms)]
 968    fig, ax = plt.subplots(figsize=(6.5, 3))
 969    for row, (name, spans, color) in enumerate([("sequential", sequential, "#c44e52"), ("parallel", parallel, "#4c72b0")]):
 970        for i, start, end in spans:
 971            ax.barh(row * 4 + i, end - start, left=start, color=color)
 972            ax.text(start + 2, row * 4 + i, f"tool {i}", va="center", fontsize=8, color="white")
 973    ax.set_yticks([1, 5], ["sequential", "parallel"])
 974    ax.invert_yaxis()
 975    ax.set_xlabel("milliseconds since the agent started the calls")
 976    ax.set_title("Three independent 100 ms tool calls")
 977    fig.tight_layout()
 978    figs["parallel"] = fig
 979    return figs
 980
 981
 982def demo() -> None:
 983    banner("1. How a call is priced (illustrative prices)")
 984    table(["call", "cost"], [
 985        ("large, 10k in / 500 out", f"${request_cost('large', 10_000, 500):.4f}"),
 986        ("same, 8k of input cached", f"${request_cost('large', 10_000, 500, cached_tokens=8_000):.4f}"),
 987        ("same, via batch interface", f"${batch_cost([('large', 10_000, 500)]):.4f}"),
 988        ("small, 10k in / 500 out", f"${request_cost('small', 10_000, 500):.4f}"),
 989    ])
 990
 991    banner("2. Routing")
 992    r = routing_savings()
 993    print(f"all large ${r['all_large']:.3f} -> routed ${r['routed']:.3f}: {r['savings']:.0%} saved")
 994    print()
 995
 996    banner("3. Semantic caching: useful and dangerous")
 997    cache = SemanticCache(threshold=0.85)
 998    cache.put("How do I reset my password?", "Use the self-service portal.")
 999    cache.put("How many vacation days do I get?", "20 days of PTO per year.")
1000    for q in ["I forgot my password, how do I recover it?", "How many sick days do I get?", "How do I reset my VPN password?"]:
1001        entry, sim = cache.lookup(q)
1002        print(f"{q!r:48} sim {sim:.2f} -> {cache.get(q)!r}")
1003    print()
1004    say("The sick-days question gets the vacation answer. That's the risk: similar is not the same.")
1005
1006    banner("4. Parallel tool calls")
1007    seq, par = asyncio.run(run_sequential([0.1] * 3)), asyncio.run(run_parallel([0.1] * 3))
1008    print(f"sequential {seq['seconds'] * 1000:.0f} ms, parallel {par['seconds'] * 1000:.0f} ms")
1009    print()
1010
1011    banner("5. Budgets and alerts")
1012    spend = TenantSpend(monthly_cap=1000)
1013    for t in [1000, 1100, 900, 1050, 950, 1000, 10_000]:
1014        alerts = spend.record("acme", 0.01, t)
1015        if alerts:
1016            print(f"task with {t} tokens -> {alerts}")
1017    print()
1018
1019    banner("6. Unit economics")
1020    table(["model", "per success (retry)", "per task (human fixes)"], [
1021        ("small $0.002 @ 60%", cost_per_success_with_retries(0.002, 0.6), cost_per_success_with_cleanup(0.002, 0.6, 2.0)),
1022        ("large $0.010 @ 95%", cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)),
1023    ])
1024
1025    banner("7. The plan, free levers first")
1026    table(["lever", "workload cost", "cumulative"], [(s["lever"], f"${s['cost']:.4f}", f"{s['factor']:.1f}x") for s in five_x_plan()])
1027    takeaway("Measure cost per successful task, take the free wins first, then route with evals watching quality.")
1028
1029
1030if __name__ == "__main__":
1031    demo()
Level 3: the code, function by function.
@dataclass(frozen=True)
class Price: on GitHub
601@dataclass(frozen=True)
602class Price:
603    input_per_m: float  # dollars per million input tokens
604    output_per_m: float  # dollars per million output tokens
Price(input_per_m: float, output_per_m: float)
input_per_m: float
output_per_m: float
PRICES: dict[str, Price] = {'small': Price(input_per_m=1.0, output_per_m=5.0), 'large': Price(input_per_m=5.0, output_per_m=25.0)}
CACHE_READ_MULTIPLIER = 0.1
BATCH_MULTIPLIER = 0.5
def request_cost( model: str, input_tokens: int, output_tokens: int, cached_tokens: int = 0) -> float: on GitHub
612def request_cost(model: str, input_tokens: int, output_tokens: int, cached_tokens: int = 0) -> float:
613    """Dollar cost of one call. `cached_tokens` of the input are billed at the cached rate."""
614    p = PRICES[model]
615    uncached = input_tokens - cached_tokens
616    return (cached_tokens * CACHE_READ_MULTIPLIER * p.input_per_m + uncached * p.input_per_m + output_tokens * p.output_per_m) / 1e6

Dollar cost of one call. cached_tokens of the input are billed at the cached rate.

def batch_cost(requests: list[tuple[str, int, int]]) -> float: on GitHub
619def batch_cost(requests: list[tuple[str, int, int]]) -> float:
620    """Cost of (model, input, output) requests submitted through a batch interface."""
621    return BATCH_MULTIPLIER * sum(request_cost(m, i, o) for m, i, o in requests)

Cost of (model, input, output) requests submitted through a batch interface.

HARD_SIGNALS = ('plan', 'analy', 'design', 'strategy', 'compare', 'debug', 'reason', 'multi-step', 'trade-off', 'why')
def route(task_text: str, max_simple_chars: int = 400) -> str: on GitHub
633def route(task_text: str, max_simple_chars: int = 400) -> str:
634    """Return "small" or "large" for a task."""
635    text = task_text.lower()
636    if len(text) > max_simple_chars or any(s in text for s in HARD_SIGNALS):
637        return "large"
638    return "small"

Return "small" or "large" for a task.

@dataclass
class WorkItem: on GitHub
641@dataclass
642class WorkItem:
643    text: str
644    output_tokens: int
645    interactive: bool = True
646    # Input is split so the levers have something to act on.
647    stable_prefix: int = 8_000  # system prompt + tool definitions, identical every call
648    tool_output: int = 3_000  # verbose tool results that could be compressed
649    rest: int = 1_000  # the question and its specific context
WorkItem( text: str, output_tokens: int, interactive: bool = True, stable_prefix: int = 8000, tool_output: int = 3000, rest: int = 1000)
text: str
interactive: bool = True
stable_prefix: int = 8000
tool_output: int = 3000
rest: int = 1000
SAMPLE_WORKLOAD: list[WorkItem] = [WorkItem(text='Classify this ticket as billing, technical or other', output_tokens=150, interactive=False, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Classify this ticket as billing, technical or other', output_tokens=150, interactive=False, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Classify this ticket as billing, technical or other', output_tokens=150, interactive=False, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Classify this ticket as billing, technical or other', output_tokens=150, interactive=False, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Extract the invoice date and total from this email', output_tokens=150, interactive=True, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Extract the invoice date and total from this email', output_tokens=150, interactive=True, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Extract the invoice date and total from this email', output_tokens=150, interactive=True, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Plan the migration of our billing system and analyze the risks of each step', output_tokens=600, interactive=True, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Plan the migration of our billing system and analyze the risks of each step', output_tokens=600, interactive=True, stable_prefix=8000, tool_output=3000, rest=1000), WorkItem(text='Plan the migration of our billing system and analyze the risks of each step', output_tokens=600, interactive=True, stable_prefix=8000, tool_output=3000, rest=1000)]
def workload_cost( work: list[WorkItem], cache: bool = False, trim: bool = False, concise: bool = False, routed: bool = False, batch: bool = False) -> float: on GitHub
659def workload_cost(work: list[WorkItem], cache: bool = False, trim: bool = False, concise: bool = False,
660                  routed: bool = False, batch: bool = False) -> float:
661    """Cost of a workload with any combination of levers switched on."""
662    total = 0.0
663    for w in work:
664        tool = 500 if trim else w.tool_output  # compress tool output to the needed fields
665        out = int(w.output_tokens * 2 / 3) if concise and w.output_tokens > 300 else w.output_tokens
666        n_in = w.stable_prefix + tool + w.rest
667        model = route(w.text) if routed else "large"
668        c = request_cost(model, n_in, out, cached_tokens=w.stable_prefix if cache else 0)
669        if batch and not w.interactive:
670            c *= BATCH_MULTIPLIER
671        total += c
672    return total

Cost of a workload with any combination of levers switched on.

def routing_savings( work: list[WorkItem] | None = None) -> dict[str, float]: on GitHub
675def routing_savings(work: list[WorkItem] | None = None) -> dict[str, float]:
676    work = work or SAMPLE_WORKLOAD
677    all_large = workload_cost(work)
678    routed = workload_cost(work, routed=True)
679    return {"all_large": all_large, "routed": routed, "savings": 1 - routed / all_large}
class ResponseCache: on GitHub
691class ResponseCache:
692    """Exact-match cache: free and never wrong, but misses every paraphrase."""
693
694    def __init__(self) -> None:
695        self._store: dict[str, str] = {}
696
697    def put(self, question: str, answer: str) -> None:
698        self._store[_normalize(question)] = answer
699
700    def get(self, question: str) -> str | None:
701        return self._store.get(_normalize(question))

Exact-match cache: free and never wrong, but misses every paraphrase.

def put(self, question: str, answer: str) -> None: on GitHub
697    def put(self, question: str, answer: str) -> None:
698        self._store[_normalize(question)] = answer
def get(self, question: str) -> str | None: on GitHub
700    def get(self, question: str) -> str | None:
701        return self._store.get(_normalize(question))
class SemanticCache: on GitHub
724class SemanticCache:
725    """Answer similar questions from memory, with guards against the obvious wrong hits.
726
727    Guards: identifiers and numbers must match exactly, negation must match,
728    and entries expire after `ttl_seconds`. None of this can separate
729    "sick days" from "vacation days"; only narrow scoping and a measured
730    wrong-hit rate can.
731    """
732
733    def __init__(self, embedder: Any = None, threshold: float = 0.9, ttl_seconds: float | None = None,
734                 clock: Callable[[], float] = time.monotonic):
735        self.embedder = embedder or ConceptEmbedder()
736        self.threshold = threshold
737        self.ttl = ttl_seconds
738        self.clock = clock
739        self.entries: list[_Entry] = []
740
741    def put(self, question: str, answer: str) -> None:
742        self.entries.append(_Entry(self.embedder.encode(question), question, answer, self.clock()))
743
744    def lookup(self, question: str) -> tuple[_Entry | None, float]:
745        """Nearest stored entry that passes the guards, and its similarity."""
746        q = self.embedder.encode(question)
747        best, best_sim = None, -1.0
748        for e in self.entries:
749            if self.ttl is not None and self.clock() - e.stored_at > self.ttl:
750                continue
751            if _identifiers(e.question) != _identifiers(question) or _negated(e.question) != _negated(question):
752                continue
753            sim = float(e.vector @ q)  # vectors are unit length, so dot product = cosine similarity
754            if sim > best_sim:
755                best, best_sim = e, sim
756        return best, best_sim
757
758    def get(self, question: str) -> str | None:
759        entry, sim = self.lookup(question)
760        return entry.answer if entry is not None and sim >= self.threshold else None

Answer similar questions from memory, with guards against the obvious wrong hits.

Guards: identifiers and numbers must match exactly, negation must match, and entries expire after ttl_seconds. None of this can separate "sick days" from "vacation days"; only narrow scoping and a measured wrong-hit rate can.

SemanticCache( embedder: Any = None, threshold: float = 0.9, ttl_seconds: float | None = None, clock: Callable[[], float] = <built-in function monotonic>) on GitHub
733    def __init__(self, embedder: Any = None, threshold: float = 0.9, ttl_seconds: float | None = None,
734                 clock: Callable[[], float] = time.monotonic):
735        self.embedder = embedder or ConceptEmbedder()
736        self.threshold = threshold
737        self.ttl = ttl_seconds
738        self.clock = clock
739        self.entries: list[_Entry] = []
embedder
threshold
ttl
clock
entries: list[primer.agents.cost._Entry]
def put(self, question: str, answer: str) -> None: on GitHub
741    def put(self, question: str, answer: str) -> None:
742        self.entries.append(_Entry(self.embedder.encode(question), question, answer, self.clock()))
def lookup(self, question: str) -> tuple[primer.agents.cost._Entry | None, float]: on GitHub
744    def lookup(self, question: str) -> tuple[_Entry | None, float]:
745        """Nearest stored entry that passes the guards, and its similarity."""
746        q = self.embedder.encode(question)
747        best, best_sim = None, -1.0
748        for e in self.entries:
749            if self.ttl is not None and self.clock() - e.stored_at > self.ttl:
750                continue
751            if _identifiers(e.question) != _identifiers(question) or _negated(e.question) != _negated(question):
752                continue
753            sim = float(e.vector @ q)  # vectors are unit length, so dot product = cosine similarity
754            if sim > best_sim:
755                best, best_sim = e, sim
756        return best, best_sim

Nearest stored entry that passes the guards, and its similarity.

def get(self, question: str) -> str | None: on GitHub
758    def get(self, question: str) -> str | None:
759        entry, sim = self.lookup(question)
760        return entry.answer if entry is not None and sim >= self.threshold else None
CACHE_PAIRS: list[tuple[str, str, bool]] = [('How do I reset my password?', 'I forgot my password, how do I recover it?', True), ('How do I reset my password?', 'How do I reset my VPN password?', False), ('How many vacation days do I get?', 'How much PTO do I get per year?', True), ('How many vacation days do I get?', 'How many sick days do I get?', False), ('How do I get reimbursed for car mileage?', 'automobile expense reimbursement', True), ('Cancel my order', "Don't cancel my order", False), ('What does ERR-4012 mean?', 'What does ERR-4013 mean?', False), ('How do I set up the VPN?', 'VPN setup steps', True), ('Report a phishing email', 'How do I report a suspicious email?', True), ('Request a new laptop', 'Request a new monitor', False), ('What is the travel per-diem?', 'How much is the daily meal per-diem when traveling?', True), ('Printer not working', 'printing fails on floor printer', True)]
def semantic_cache_sweep(thresholds: list[float]) -> list[dict[str, float]]: on GitHub
780def semantic_cache_sweep(thresholds: list[float]) -> list[dict[str, float]]:
781    """For each threshold: share of paraphrases served from cache, share of near-misses served wrongly."""
782    out = []
783    for th in thresholds:
784        useful = wrong = 0
785        for stored, new, same in CACHE_PAIRS:
786            cache = SemanticCache(threshold=th)
787            cache.put(stored, "cached answer")
788            hit = cache.get(new) is not None
789            useful += hit and same
790            wrong += hit and not same
791        n_same = sum(1 for p in CACHE_PAIRS if p[2])
792        out.append({"threshold": th, "useful_hit_rate": useful / n_same, "wrong_hit_rate": wrong / (len(CACHE_PAIRS) - n_same)})
793    return out

For each threshold: share of paraphrases served from cache, share of near-misses served wrongly.

async def run_sequential(delays: list[float]) -> dict[str, typing.Any]: on GitHub
808async def run_sequential(delays: list[float]) -> dict[str, Any]:
809    t0, log = time.perf_counter(), []
810    results = [await _fake_tool(i, d, t0, log) for i, d in enumerate(delays)]
811    return {"seconds": time.perf_counter() - t0, "results": results, "spans": sorted(log)}
async def run_parallel(delays: list[float]) -> dict[str, typing.Any]: on GitHub
814async def run_parallel(delays: list[float]) -> dict[str, Any]:
815    t0, log = time.perf_counter(), []
816    # gather starts every call before awaiting any, and returns results in call order.
817    results = await asyncio.gather(*(_fake_tool(i, d, t0, log) for i, d in enumerate(delays)))
818    return {"seconds": time.perf_counter() - t0, "results": list(results), "spans": sorted(log)}
class BudgetExceeded(builtins.RuntimeError): on GitHub
826class BudgetExceeded(RuntimeError):
827    pass

Unspecified run-time error.

@dataclass
class TaskBudget: on GitHub
830@dataclass
831class TaskBudget:
832    """Hard per-task limits. The agent loop calls charge() before every model call."""
833
834    max_steps: int
835    max_tokens: int
836    steps: int = 0
837    tokens: int = 0
838
839    def charge(self, tokens: int) -> None:
840        if self.steps + 1 > self.max_steps:
841            raise BudgetExceeded(f"step limit {self.max_steps} reached")
842        if self.tokens + tokens > self.max_tokens:
843            raise BudgetExceeded(f"token limit {self.max_tokens} would be exceeded ({self.tokens} + {tokens})")
844        self.steps += 1
845        self.tokens += tokens

Hard per-task limits. The agent loop calls charge() before every model call.

TaskBudget(max_steps: int, max_tokens: int, steps: int = 0, tokens: int = 0)
max_steps: int
max_tokens: int
steps: int = 0
tokens: int = 0
def charge(self, tokens: int) -> None: on GitHub
839    def charge(self, tokens: int) -> None:
840        if self.steps + 1 > self.max_steps:
841            raise BudgetExceeded(f"step limit {self.max_steps} reached")
842        if self.tokens + tokens > self.max_tokens:
843            raise BudgetExceeded(f"token limit {self.max_tokens} would be exceeded ({self.tokens} + {tokens})")
844        self.steps += 1
845        self.tokens += tokens
@dataclass
class TenantSpend: on GitHub
848@dataclass
849class TenantSpend:
850    """Per-tenant monthly spend cap plus an anomaly alert on tokens per task."""
851
852    monthly_cap: float
853    z: float = 3.0
854    min_history: int = 5
855    spent: dict[str, float] = field(default_factory=dict)
856    history: dict[str, list[int]] = field(default_factory=dict)
857
858    def record(self, tenant: str, cost: float, tokens: int) -> dict[str, str]:
859        """Record one finished task; return alerts keyed by kind ("cap", "anomaly")."""
860        alerts: dict[str, str] = {}
861        self.spent[tenant] = self.spent.get(tenant, 0.0) + cost
862        if self.spent[tenant] > self.monthly_cap:
863            alerts["cap"] = f"{tenant} spent {self.spent[tenant]:.2f} of a {self.monthly_cap:.2f} cap"
864        hist = self.history.setdefault(tenant, [])
865        if len(hist) >= self.min_history:
866            mu, sigma = statistics.mean(hist), statistics.stdev(hist)
867            if tokens > mu + self.z * max(sigma, 1.0):
868                alerts["anomaly"] = f"{tokens} tokens vs usual {mu:.0f} ± {sigma:.0f}"
869        hist.append(tokens)
870        return alerts

Per-tenant monthly spend cap plus an anomaly alert on tokens per task.

TenantSpend( monthly_cap: float, z: float = 3.0, min_history: int = 5, spent: dict[str, float] = <factory>, history: dict[str, list[int]] = <factory>)
monthly_cap: float
z: float = 3.0
min_history: int = 5
spent: dict[str, float]
history: dict[str, list[int]]
def record(self, tenant: str, cost: float, tokens: int) -> dict[str, str]: on GitHub
858    def record(self, tenant: str, cost: float, tokens: int) -> dict[str, str]:
859        """Record one finished task; return alerts keyed by kind ("cap", "anomaly")."""
860        alerts: dict[str, str] = {}
861        self.spent[tenant] = self.spent.get(tenant, 0.0) + cost
862        if self.spent[tenant] > self.monthly_cap:
863            alerts["cap"] = f"{tenant} spent {self.spent[tenant]:.2f} of a {self.monthly_cap:.2f} cap"
864        hist = self.history.setdefault(tenant, [])
865        if len(hist) >= self.min_history:
866            mu, sigma = statistics.mean(hist), statistics.stdev(hist)
867            if tokens > mu + self.z * max(sigma, 1.0):
868                alerts["anomaly"] = f"{tokens} tokens vs usual {mu:.0f} ± {sigma:.0f}"
869        hist.append(tokens)
870        return alerts

Record one finished task; return alerts keyed by kind ("cap", "anomaly").

def cost_per_success_with_retries(cost_per_attempt: float, success_rate: float) -> float: on GitHub
878def cost_per_success_with_retries(cost_per_attempt: float, success_rate: float) -> float:
879    """Retry until success: expected attempts are 1/p."""
880    return cost_per_attempt / success_rate

Retry until success: expected attempts are 1/p.

def cost_per_success_with_cleanup( cost_per_attempt: float, success_rate: float, human_fix_cost: float) -> float: on GitHub
883def cost_per_success_with_cleanup(cost_per_attempt: float, success_rate: float, human_fix_cost: float) -> float:
884    """One attempt per task; a person fixes each failure."""
885    return cost_per_attempt + (1 - success_rate) * human_fix_cost

One attempt per task; a person fixes each failure.

def five_x_plan( work: list[WorkItem] | None = None) -> list[dict[str, typing.Any]]: on GitHub
888def five_x_plan(work: list[WorkItem] | None = None) -> list[dict[str, Any]]:
889    """Apply the levers cumulatively, free ones first. Returns cost after each."""
890    work = work or SAMPLE_WORKLOAD
891    levers = [
892        ("baseline: everything on the large model", {}),
893        ("prompt caching", {"cache": True}),
894        ("trim tool output", {"cache": True, "trim": True}),
895        ("concise output", {"cache": True, "trim": True, "concise": True}),
896        ("route easy tasks to the small model", {"cache": True, "trim": True, "concise": True, "routed": True}),
897        ("batch non-interactive work", {"cache": True, "trim": True, "concise": True, "routed": True, "batch": True}),
898    ]
899    base = workload_cost(work)
900    steps = []
901    for name, kw in levers:
902        c = workload_cost(work, **kw)
903        steps.append({"lever": name, "cost": c, "factor": base / c})
904    return steps

Apply the levers cumulatively, free ones first. Returns cost after each.

def figures() -> dict[str, typing.Any]: on GitHub
912def figures() -> dict[str, Any]:
913    import matplotlib
914
915    matplotlib.use("Agg")
916    import matplotlib.pyplot as plt
917
918    figs: dict[str, Any] = {}
919
920    # 1. Unit economics.
921    fig, ax = plt.subplots(figsize=(6.5, 3.5))
922    labels = ["retry until success", "person fixes failures"]
923    small = [cost_per_success_with_retries(0.002, 0.60), cost_per_success_with_cleanup(0.002, 0.60, 2.0)]
924    large = [cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)]
925    x = np.arange(2)
926    ax.bar(x - 0.2, small, 0.4, label="small: $0.002/attempt, 60% success", color="#c44e52")
927    ax.bar(x + 0.2, large, 0.4, label="large: $0.010/attempt, 95% success", color="#4c72b0")
928    for xi, (s, lg) in enumerate(zip(small, large)):
929        ax.text(xi - 0.2, s, f"${s:.3f}", ha="center", va="bottom", fontsize=8)
930        ax.text(xi + 0.2, lg, f"${lg:.3f}", ha="center", va="bottom", fontsize=8)
931    ax.set_yscale("log")
932    ax.set_xticks(x, labels)
933    ax.set_ylabel("dollars per successful task (log scale)")
934    ax.set_title("The cheap model is only cheap if failures are free")
935    ax.legend(fontsize=8, loc="upper left")
936    fig.tight_layout()
937    figs["unit_economics"] = fig
938
939    # 2. The 5x plan.
940    steps = five_x_plan()
941    fig, ax = plt.subplots(figsize=(7.5, 3.8))
942    ax.bar(range(len(steps)), [s["cost"] for s in steps], color=["#8c8c8c"] + ["#4c72b0"] * 3 + ["#dd8452"] * 2)
943    for i, s in enumerate(steps):
944        ax.text(i, s["cost"], f"{s['factor']:.1f}x", ha="center", va="bottom", fontsize=9)
945    ax.set_xticks(range(len(steps)), ["baseline", "+ caching", "+ trim", "+ concise", "+ routing", "+ batch"])
946    ax.set_ylabel("workload cost ($)")
947    ax.set_title("Levers applied in order (blue: can't change quality; orange: validate with evals)")
948    fig.tight_layout()
949    figs["five_x"] = fig
950
951    # 3. Semantic cache threshold sweep.
952    ths = list(np.round(np.arange(0.3, 1.001, 0.025), 3))
953    sweep = semantic_cache_sweep(ths)
954    fig, ax = plt.subplots(figsize=(6.5, 3.5))
955    ax.plot(ths, [s["useful_hit_rate"] for s in sweep], color="#4c72b0", label="paraphrases served from cache (good)")
956    ax.plot(ths, [s["wrong_hit_rate"] for s in sweep], color="#c44e52", label="different questions served from cache (wrong)")
957    ax.set_xlabel("similarity threshold")
958    ax.set_ylabel("share of pairs")
959    ax.set_title("Semantic cache: no threshold separates them cleanly")
960    ax.legend(fontsize=8)
961    fig.tight_layout()
962    figs["semantic_cache"] = fig
963
964    # 4. Parallel vs sequential timeline: the schedule each approach follows. Drawn from the delays, not
965    # a stopwatch, so the site doesn't change with a millisecond of jitter; the demo measures the real thing.
966    delays_ms = [100, 100, 100]
967    sequential = [(i, sum(delays_ms[:i]), sum(delays_ms[: i + 1])) for i in range(len(delays_ms))]
968    parallel = [(i, 0, d) for i, d in enumerate(delays_ms)]
969    fig, ax = plt.subplots(figsize=(6.5, 3))
970    for row, (name, spans, color) in enumerate([("sequential", sequential, "#c44e52"), ("parallel", parallel, "#4c72b0")]):
971        for i, start, end in spans:
972            ax.barh(row * 4 + i, end - start, left=start, color=color)
973            ax.text(start + 2, row * 4 + i, f"tool {i}", va="center", fontsize=8, color="white")
974    ax.set_yticks([1, 5], ["sequential", "parallel"])
975    ax.invert_yaxis()
976    ax.set_xlabel("milliseconds since the agent started the calls")
977    ax.set_title("Three independent 100 ms tool calls")
978    fig.tight_layout()
979    figs["parallel"] = fig
980    return figs
def demo() -> None: on GitHub
 983def demo() -> None:
 984    banner("1. How a call is priced (illustrative prices)")
 985    table(["call", "cost"], [
 986        ("large, 10k in / 500 out", f"${request_cost('large', 10_000, 500):.4f}"),
 987        ("same, 8k of input cached", f"${request_cost('large', 10_000, 500, cached_tokens=8_000):.4f}"),
 988        ("same, via batch interface", f"${batch_cost([('large', 10_000, 500)]):.4f}"),
 989        ("small, 10k in / 500 out", f"${request_cost('small', 10_000, 500):.4f}"),
 990    ])
 991
 992    banner("2. Routing")
 993    r = routing_savings()
 994    print(f"all large ${r['all_large']:.3f} -> routed ${r['routed']:.3f}: {r['savings']:.0%} saved")
 995    print()
 996
 997    banner("3. Semantic caching: useful and dangerous")
 998    cache = SemanticCache(threshold=0.85)
 999    cache.put("How do I reset my password?", "Use the self-service portal.")
1000    cache.put("How many vacation days do I get?", "20 days of PTO per year.")
1001    for q in ["I forgot my password, how do I recover it?", "How many sick days do I get?", "How do I reset my VPN password?"]:
1002        entry, sim = cache.lookup(q)
1003        print(f"{q!r:48} sim {sim:.2f} -> {cache.get(q)!r}")
1004    print()
1005    say("The sick-days question gets the vacation answer. That's the risk: similar is not the same.")
1006
1007    banner("4. Parallel tool calls")
1008    seq, par = asyncio.run(run_sequential([0.1] * 3)), asyncio.run(run_parallel([0.1] * 3))
1009    print(f"sequential {seq['seconds'] * 1000:.0f} ms, parallel {par['seconds'] * 1000:.0f} ms")
1010    print()
1011
1012    banner("5. Budgets and alerts")
1013    spend = TenantSpend(monthly_cap=1000)
1014    for t in [1000, 1100, 900, 1050, 950, 1000, 10_000]:
1015        alerts = spend.record("acme", 0.01, t)
1016        if alerts:
1017            print(f"task with {t} tokens -> {alerts}")
1018    print()
1019
1020    banner("6. Unit economics")
1021    table(["model", "per success (retry)", "per task (human fixes)"], [
1022        ("small $0.002 @ 60%", cost_per_success_with_retries(0.002, 0.6), cost_per_success_with_cleanup(0.002, 0.6, 2.0)),
1023        ("large $0.010 @ 95%", cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)),
1024    ])
1025
1026    banner("7. The plan, free levers first")
1027    table(["lever", "workload cost", "cumulative"], [(s["lever"], f"${s['cost']:.4f}", f"{s['factor']:.1f}x") for s in five_x_plan()])
1028    takeaway("Measure cost per successful task, take the free wins first, then route with evals watching quality.")