primer.agents.cost
Cost and latency: paying less per successful task without getting worse
Run: python -m primer.agents.cost
This lesson builds on tokens and pricing from primer.ml.inference, on the
context window from primer.agents.context and on the release gate from
primer.agents.evals.
Level 1: The practitioner's guide
In one sentence. Cost engineering for a model-backed system means paying less per successful task, not per call, by taking the levers that can't hurt quality first and checking the ones that can against an eval.
When you need it. When the bill grows faster than the value, when a latency budget is missed, or before the first big customer arrives and the price per task stops being a rounding error. The tell: you know your spend per month but not your cost per successful task, or you know cost per call but have never counted retries and human clean-up. The lesson's worked example shows why that number, and only that number, decides things: a small model at \$0.002 a call with a 60% success rate beats a large one at \$0.010 and 95% if failures can simply be retried (\$0.0033 against \$0.0105 per success), and loses by 7x if a person has to fix each failure (\$0.802 against \$0.110 per task, at an illustrative \$2 per fix). You don't need any of this for a prototype with ten users a day, and you shouldn't touch the levers that change behaviour (routing, semantic caching) until you have an eval to watch them with. All prices in this lesson are illustrative constants chosen for round arithmetic (a large model at \$5 per million input tokens and \$25 per million output, a small one at \$1 and \$5); check a provider's current price list for real numbers.
Your options. Eight levers, from the ones that can't hurt quality to the ones that can:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Prompt caching | Puts the stable part of the prompt first so the provider reuses its work | No behaviour change; the worked 10,000-token call falls from \$0.0625 to \$0.0265 with 8,000 tokens cached | A small write premium the first time; reordering the prompt | The provider's API |
| Trim tokens | Compress tool results, send fewer retrieved chunks, ask for concise output | No change if you only cut what the next step never read; output is the expensive direction | Engineering, and care not to cut what was needed | Your prompt and code |
| Parallel tool calls | Run independent tool calls at the same time | Wall-clock time falls to the slowest call (300 ms to 100 ms for three lookups); tokens unchanged | Concurrency in the agent loop | Your agent loop |
| Exact response cache | The same question, ignoring case and spaces, returns the stored answer | Never wrong until the facts change | A time-to-live to manage | Your code |
| Batch API | Non-interactive work submitted in bulk at about half price | Half price, results within hours | Waiting; only for work nobody is waiting on | The provider's API |
| Budgets and alerts | Caps steps and tokens per task and spend per tenant; alerts on a sudden jump | A confused agent stops instead of looping all night | Choosing the limits; a false alarm now and then | Your code |
| Routing | Sends easy tasks (classify, extract, format) to a small model and hard ones to a large one | 53% saved on this lesson's ten-task workload, if the router is right | The first lever that can lose quality; needs an eval before and after | Your code, in front of every request |
| Semantic cache | Answers a similar question from memory, using embeddings to judge similarity | The cheapest possible hit | Confidently wrong answers: on this lesson's pairs, 40% of different-intent questions are served from cache at a 0.95 threshold | Your code, plus an embedding model |
A ninth, for later: train or distil a small model on the large model's answers so it can take more of the routed traffic (Hinton et al., 2015).
How to choose. Measure first, then go down the table in order.
- No measurement yet: instrument cost per task by step, by model, input against output and cached against uncached, and put an eval set in place to hold quality fixed.
- A long, stable system prompt or tool list: prompt caching, today. It is the only lever that halves a bill without changing a single answer.
- Big tool results or many retrieved chunks: trim. Compress to the fields the next step needs and send the top few reranked chunks.
- Several independent lookups per turn: run them concurrently, and stream the answer so users see progress.
- Anything nobody is waiting on (nightly evals, backfills, bulk classification): batch it.
- Most traffic is simple: route, with the eval watching. A router that sends hard tasks to the small model saves money and quietly loses quality.
- Narrow, curated, FAQ-style traffic and nothing else: a semantic cache with guards (identifiers, numbers and negation must match), a time-to-live and a measured wrong-hit rate. On everyday questions no threshold makes it safe.
- Whatever you pick, judge it by cost per successful task, with retries and human clean-up counted in.
What it costs. The levers stack. On this lesson's ten-task workload the baseline is \$0.671; caching takes it to \$0.311 (2.2x), trimming tool output to \$0.186 (3.6x), concise output to \$0.171 (3.9x), routing to \$0.086 (7.8x) and batching the non-interactive share to \$0.080 (8.4x). The first three are free wins that don't touch quality; routing is the big one and the first that can. Caching is not free on the first request: Anthropic's prompt caching documentation, for one, prices a cache write at 1.25 times the base input price for a five-minute cache (2 times for an hour), reads at a tenth of it or less, and only caches prompts above a per-model minimum of a few hundred to a few thousand tokens. Batching costs time: the same provider quotes a 50% discount with most batches finishing within an hour. Parallel calls and streaming cost no tokens at all; they buy time and perceived speed. Budgets cost the occasional false alarm: this lesson's anomaly rule flags a task using more than the mean plus three standard deviations of recent usage, so a 10,000-token task against a history around 1,000 trips it at once.
What breaks.
- Cost per call replacing cost per success. The cheap model looks cheaper until failures are priced. Count retries and fixes.
- Routing that loses quality quietly. Nothing errors; the answers just get worse. Gate the router with the eval, and re-check when the small model changes.
- The semantic cache that answers the wrong question. "How many sick days do I get?" scores 0.97 against the vacation question and gets the vacation answer. Near-misses score higher than real paraphrases, so no threshold separates them.
- Stale cache hits. The policy changed; the cache didn't. Give every entry a time-to-live.
- A cache that never hits. Something volatile (a timestamp, the user's name) sits at the top of the prompt, so the prefix differs every call. Stable content first.
- Trimming what was needed. A tool result cut to 500 tokens loses the field the next step reads. Compress by field, not by length alone.
- The runaway loop. An agent calls the same tool with the same arguments all night. Cap steps and tokens per task; alert on tokens per task jumping.
- One tenant's surprise invoice. A shared platform with no per-customer cap turns one customer's bug into everyone's bill.
In the wild. Prompt caching and batch processing are standard features
of hosted APIs; Anthropic's documentation for both is linked in Further
reading and gives the multipliers quoted above. FrugalGPT (Chen, Zaharia and
Zou, 2023) named the three families, prompt adaptation, model approximation
and cascades that try a cheap model first and escalate, and reported
matching the best single model with up to 98% less cost on its benchmarks.
Distillation (Hinton, Vinyals and Dean, 2015) is how the small model in the
routing table gets good enough to take more traffic. Tooling: LiteLLM is an
open-source gateway that puts one interface in front of many providers and
adds routing with fallbacks, per-key and per-team budgets, spend tracking
and caching; GPTCache is an open-source semantic cache built on embeddings
and a vector store with a pluggable similarity evaluator, the design this
lesson's SemanticCache reproduces in miniature, guards and all. Parallel
tool use is part of the tool-calling protocol of the major APIs: the model
asks for several tools in one turn and you return all the results in one
message.
Go deeper. Level 2 prices a request symbol by symbol, builds the router, both caches and the parallel loop in plain Python, derives the anomaly alert from a mean and a standard deviation, works the unit economics with every number, and stacks the levers to draw the 8.4x figure. If you only needed to know which lever to pull first, you are done.
Level 2: How it works, from scratch
A restaurant doesn't cut costs by buying worse ingredients for every dish. It sends simple orders to the line cook and complex ones to the head chef, pre-chops what every dish shares, doesn't plate food nobody eats, cooks things in parallel, does bulk prep overnight when it's cheaper, and watches the one kitchen station whose bills suddenly spike.
Every lever in this lesson is one of those moves. The measure that matters is cost per successful task, not cost per call: a cheap attempt that fails still has to be paid for, and so does fixing it.
All prices in this module are illustrative constants chosen for round arithmetic (a "large" model at \$5 per million input tokens and \$25 per million output tokens; a "small" one at \$1 and \$5). They show the shape of the trade-offs; check your provider's current price list for real numbers.
How a request is priced
Everyday picture. A taxi that charges one rate for the distance to your
pickup and a higher rate for the ride itself. Models charge per token
(a word piece, about 4 characters of English): one rate for tokens you send
(input) and a higher rate for tokens the model writes (output),
because writing happens one token at a time (see primer.ml.inference).
Worked example. A large-model call with 10,000 input tokens and 500 output tokens: 10,000 × \$5 / 1,000,000 = \$0.05 for input, plus 500 × \$25 / 1,000,000 = \$0.0125 for output, total \$0.0625. If 8,000 of those input tokens are a stable prefix served from the prompt cache (reused work from an earlier identical beginning, billed here at 10% of the input price), input becomes 8,000 × \$0.50/M + 2,000 × \$5/M = \$0.014, and the call costs \$0.0265, 58% less.
Level 3: the formula and its symbols
$$ \text{cost} = \frac{c \cdot \rho\, p_{\text{in}} + (n_{\text{in}} - c)\, p_{\text{in}} + n_{\text{out}}\, p_{\text{out}}}{10^6} $$
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| $n_{\text{in}}$ | input tokens sent | 10,000 |
| $c$ | input tokens served from the prompt cache | 8,000 |
| $n_{\text{out}}$ | output tokens generated | 500 |
| $p_{\text{in}}, p_{\text{out}}$ | price per million input / output tokens | \$5, \$25 |
| $\rho$ | cached-read price as a fraction of normal input price (Greek rho) | 0.1 |
| $10^6$ | prices are per million tokens |
In words: cached input at its discounted rate, plus the rest of the input at full rate, plus output at the output rate, all per million.
On the worked example: (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ = (4,000 + 10,000 + 12,500) / 10⁶ = \$0.0265.
Level 3: in Python
In Python:
n_in, c, n_out = 10_000, 8_000, 500
p_in, p_out, rho = 5, 25, 0.1
cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6
round(cost, 4) # → 0.0265
Two facts fall out: output tokens cost several times more than input (ask
for concise answers), and a cached prefix is nearly free (put stable content
first; see primer.agents.context). Caches usually charge a small premium
the first time a prefix is written; this module ignores it for simplicity.
In code: request_cost is the formula above, reading each model's
Price from PRICES.
Route each task to the cheapest model that can do it
Everyday picture. A hospital triage nurse: sprained ankles go to the nurse practitioner, chest pains to the cardiologist. Nobody sends every patient to the most expensive specialist.
Worked example. The sample workload is 10 tasks: 7 simple (classify a ticket, extract a date) and 3 complex (plan a migration). All on the large model it costs \$0.671. Routing the 7 simple ones to the small model, which is 5x cheaper per token, brings it to \$0.314: 53% saved, with the hard tasks still on the strong model.
flowchart LR T[Incoming task] --> R{Router<br/>rules or a small classifier} R -->|classify, extract,<br/>format, short| S[Small, fast model] R -->|plan, analyze,<br/>multi-step reasoning| L[Large model] S --> O[Result] L --> O
Reading it: the router is cheap (a few rules here; often a small
classifier) and sits in front of every request. The whole saving comes from
the left branch: most traffic in real systems is simple, and simple work
runs well on small models. Validate the router with evals
(primer.agents.evals): a router that sends hard tasks to the small model
saves money and quietly loses quality.
In code: route is the rule-based router; workload_cost prices a list
of WorkItem tasks with any levers switched on; routing_savings compares
all-large against routed on SAMPLE_WORKLOAD.
Response caching and semantic caching
Everyday picture. A receptionist who's been asked "what's the wifi
password?" a hundred times just answers from memory. That's an exact
response cache: the same question (ignoring case and spaces) returns the
stored answer with no model call. A semantic cache goes further and
answers similar questions from memory, using embeddings (vectors where
closeness means similar meaning, see primer.ml.embeddings) to decide
"similar". That's where the receptionist starts giving the vacation-policy
answer to someone asking about sick leave.
Worked example. With the toy embedder in primer.common.embedder:
| Stored question | New question | Similarity | Same intent? |
|---|---|---|---|
| How do I reset my password? | I forgot my password, how do I recover it? | 0.88 | yes |
| How many vacation days do I get? | How many sick days do I get? | 0.97 | no |
| Request a new laptop | Request a new monitor | 0.96 | no |
| What does ERR-4012 mean? | What does ERR-4013 mean? | 0.56 | no |
The two near-misses score higher than a genuine paraphrase. No single threshold separates them. Two cheap guards help: identifiers and numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match ("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which share a topic and differ in a single word.
flowchart TD Q[New question] --> E{Exact match<br/>in response cache?} E -->|yes| A1[Return stored answer] E -->|no| V[Embed and find the<br/>nearest stored question] V --> T{Similarity above<br/>threshold?} T -->|no| M[Call the model] T -->|yes| G{IDs, numbers and<br/>negation identical?<br/>entry not expired?} G -->|no| M G -->|yes| A2[Return stored answer] M --> S[Store the new answer]
Reading it: the exact cache is checked first because it's free and never wrong. The semantic path adds two gates: the similarity threshold, and the guards. An expired entry (older than its TTL, time to live) is treated as a miss, so answers about changing facts don't go stale forever.
Reading it: the x-axis is the similarity threshold; the blue line is the share of genuine paraphrases answered from cache (good), and the red line is the share of different-intent questions answered from cache (a wrong answer served confidently). Lowering the threshold raises both, but look where the lines sit: from 0.4 upward the red line is above the blue one, so the cache serves more wrong answers than right ones (60% of different-intent questions against at most 57% of paraphrases). Even at 0.95 it still answers 40% of them wrongly and only 14% of paraphrases, because "sick" vs "vacation" scores 0.97. On everyday questions no threshold makes semantic caching safe; it's safe only for narrow, curated FAQ-style traffic, with guards, a TTL, and a measured wrong-hit rate.
In code: ResponseCache is the exact cache. SemanticCache is the
semantic path: SemanticCache.lookup skips expired entries and entries whose
identifiers or negation differ, then returns the nearest survivor, and
SemanticCache.get applies the threshold. semantic_cache_sweep draws the
figure from the labelled pairs in CACHE_PAIRS.
Trim tokens
Everyday picture. Don't photocopy the whole binder when the colleague
needs one page. Compress tool results to the fields the next step needs
(primer.agents.context.compress_tool_output), send the top few reranked
chunks instead of dozens, remove repeated boilerplate from system prompts,
and ask for concise output, because output is the expensive direction.
In code: workload_cost models trimming as cutting each task's tool
output to 500 tokens, and concise output as cutting long answers by a third.
Run independent tool calls at the same time
Everyday picture. Boil the pasta while the sauce simmers. Cooking them one after the other takes the sum of the times; together, the time of the slowest.
Worked example. Three independent lookups of 100 ms each: sequentially about 300 ms; concurrently about 100 ms.
sequenceDiagram participant A as Agent participant T1 as Weather API participant T2 as Calendar API participant T3 as CRM API Note over A,T3: Sequential: about 300 ms A->>T1: call T1-->>A: result A->>T2: call T2-->>A: result A->>T3: call T3-->>A: result Note over A,T3: Parallel: about 100 ms par A->>T1: call and A->>T2: call and A->>T3: call end T1-->>A: result T2-->>A: result T3-->>A: result
Reading it: in the top half each call waits for the previous result; in the bottom half the three requests leave together and the agent waits only for the slowest. Models can ask for several tools in one turn (parallel tool use); run them concurrently and return all results in one message. Parallel calls cut wall-clock time, not tokens.
Reading it: each bar is one 100 ms tool call on a shared time axis. The
sequential calls stack end to end, 300 ms in all; the parallel ones start
together, so the whole batch finishes in the time of one. Run the lesson to
measure it with asyncio: the real timings land within a few milliseconds
of these bars.
In code: run_sequential awaits each call before starting the next;
run_parallel starts them all with Python's asyncio gather and waits once.
Streaming (showing tokens as they're generated) doesn't reduce total time either, but users see progress immediately, which changes how fast the system feels.
Batch what isn't interactive
Everyday picture. Sending the laundry out to be done overnight at half price instead of waiting at the express counter. Providers offer batch APIs: submit many requests, get results within hours (often within 24), typically at about half the price. Use them for anything no one is waiting on: nightly evals, backfilling document processing, bulk classification.
In code: batch_cost prices a list of requests at the batch discount.
Budgets and alerts
Everyday picture. A prepaid card with a hard limit, and a bank that texts you when a purchase looks nothing like your usual spending.
A task budget caps steps and tokens per task, so a confused agent stops instead of looping all night. A tenant spend cap (a tenant is one customer organisation on a shared platform) stops one customer's runaway usage from becoming a surprise invoice. An anomaly alert fires when a task uses far more tokens than usual:
Level 3: the formula and its symbols
$$ \text{alert if } x > \mu + z\,\sigma $$
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| $x$ | tokens used by the task just finished | 10,000 |
| $\mu$ | mean tokens per task over this tenant's recent history (Greek mu) | 1,000 |
| $\sigma$ | standard deviation of that history: the typical distance from the mean (Greek sigma) | ≈ 71 |
| $z$ | how many standard deviations count as unusual | 3 |
In words: alert when a task uses more tokens than the usual amount plus three times the usual spread.
On the worked example: history 1000, 1100, 900, 1050, 950, 1000 has mean 1,000 and standard deviation ≈ 71 (the squared distances from the mean add up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7), so the line is 1,000 + 3 × 70.7 ≈ 1,212. A 10,000-token task is far above it: alert. Such jumps usually mean a loop or a bad deploy.
Level 3: in Python
In Python:
import statistics
history = [1000, 1100, 900, 1050, 950, 1000]
mu = statistics.mean(history)
# divides by 6 - 1, as above
sigma = statistics.stdev(history)
mu, round(sigma, 1) # → (1000, 70.7)
z = 3
# the alert line
round(mu + z * sigma) # → 1212
# alert?
10_000 > mu + z * sigma # → True
In code: TaskBudget.charge counts each step's tokens and raises
BudgetExceeded at either limit; TenantSpend.record adds a finished task's
cost to its tenant's total and returns a cap alert or an anomaly alert (the
formula above).
Unit economics: cost per successful task
Everyday picture. A cheap printer that jams on 40% of pages isn't cheap once you count the wasted paper and the time spent clearing jams.
Level 3: the formula and its symbols
$$ \text{cost per success (retry until it works)} = \frac{c}{p} \qquad \text{cost per task (a person fixes failures)} = c + (1 - p)\,h $$
Symbols
| Symbol | Meaning here | Small model | Large model |
|---|---|---|---|
| $c$ | model cost of one attempt | \$0.002 | \$0.010 |
| $p$ | chance an attempt succeeds | 0.60 | 0.95 |
| $1/p$ | expected number of attempts until one succeeds | 1.67 | 1.05 |
| $h$ | cost of a person fixing one failure (illustrative: a few minutes of staff time) | \$2.00 | \$2.00 |
In words: if you can simply retry, you pay one attempt's cost for every expected attempt; if failures need a person, every task pays its attempt plus the chance of failure times the cost of the fix.
On the worked example: retrying, the small model costs 0.002 / 0.60 = \$0.0033 per success and the large one 0.010 / 0.95 = \$0.0105, so the small model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 = \$0.802 per task and the large one 0.010 + 0.05 × 2.00 = \$0.110: the "cheap" model is 7x more expensive.
Level 3: in Python
In Python:
def per_success(c, p):
# retry until it works
return c / p
def per_task(c, p, h=2.00):
# a person fixes each failure
return c + (1 - p) * h
round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4) # → (0.0033, 0.0105)
round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3) # → (0.802, 0.11)
# small vs. large, with cleanup
round(per_task(0.002, 0.60) / per_task(0.010, 0.95)) # → 7
Reading it: the left pair of bars assumes failures can be retried automatically: the small model wins. The right pair assumes a person has to fix each failure: the large model wins by a mile. Which world you're in decides which model is cheaper, and the model's price per token barely matters in the second one.
In code: cost_per_success_with_retries is $c/p$ and
cost_per_success_with_cleanup is $c + (1-p)\,h$.
Putting it together: a 5x plan, in order
Order the levers so the ones that can't hurt quality come first:
- Prompt caching: reorder the prompt so the stable prefix is reused. No behaviour change.
- Trim tokens: compress tool results and send fewer chunks.
- Concise output: ask for shorter answers where length adds nothing.
- Route easy tasks to the small model: the first lever that can change quality, so it's validated with evals before and after.
- Batch the non-interactive share.
Reading it: each bar is the cost of the same 10-task workload after applying every lever up to and including that one; the label is the cumulative reduction. Caching alone roughly halves it, trimming and concise output take it past 3x, routing takes it past 7x, and batching adds the last few percent. The three levers after the baseline (caching, trimming, concise output) are the free wins that don't touch quality.
In code: five_x_plan switches the levers on one at a time, in this
order, and reports the cost and cumulative reduction after each.
In 20 seconds
- Measure cost per successful task, not per call; include retries and human cleanup.
- Free wins first: prompt caching (stable prefix first), trimming tool output and chunks, concise output, batch APIs for non-interactive work.
- Then route easy tasks to small models, validated with evals.
- Run independent tool calls in parallel; stream to improve perceived speed.
- Budgets per task, spend caps per tenant, and alerts on sudden jumps in tokens per task.
- Semantic caches can serve confidently wrong answers; use them only on narrow traffic, with guards and a measured wrong-hit rate.
Self-test questions
How would you cut the cost per task by 5x without hurting quality? First measure: cost per successful task, broken down by step, model, input vs. output, and cached vs. uncached tokens, with an eval set to hold quality fixed. Then, in order: restructure prompts for prompt caching; compress tool results and send only the top reranked chunks; ask for concise output; route simple steps (classification, extraction, formatting) to a small model, checking the eval before and after; move anything non-interactive to a batch API; cap steps and tokens per task. On the sample workload here that's about 8x, and every step after the first three is checked against the eval set.
Why can a cheaper model cost more? Because failures cost money too: retries, human cleanup, lost customers. At 60% success with a \$2 human fix, a \$0.002 call costs \$0.80 per task; a \$0.010 call at 95% costs \$0.11.
What are the risks of a semantic cache? Returning a confident, wrong answer to a question that's similar but not the same ("sick days" vs "vacation days"), and serving stale answers after facts change. Mitigate with high thresholds, exact-match guards on IDs, numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit rate on labelled pairs before enabling it.
Your average tokens per task doubled overnight. What do you check?
Whether a deploy changed a prompt or a tool (a bigger tool output, a new
retrieval setting), whether an agent is looping (the same tool called with
the same arguments), and whether the prompt cache hit rate dropped
(something volatile moved to the top). Traces make this a lookup rather
than a guess (primer.agents.observability).
The papers behind this lesson
- Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015), https://arxiv.org/abs/1503.02531. Showed how to train a small model to imitate a large one, the technique behind making the cheap model good enough to take more of the routed traffic. annotated companion
- Chen, Zaharia & Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (2023), https://arxiv.org/abs/2305.05176. Studied prompt adaptation, caching and cascades that try cheap models first and escalate to expensive ones only when needed.
Further reading
- Anthropic prompt caching docs: https://docs.claude.com/en/docs/build-with-claude/prompt-caching
- Anthropic Message Batches docs: https://docs.claude.com/en/docs/build-with-claude/batch-processing
- Anthropic, parallel tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Python
asyncio.gather: https://docs.python.org/3/library/asyncio-task.html#asyncio.gather
1r""" 2# Cost and latency: paying less per successful task without getting worse 3 4Run: `python -m primer.agents.cost` 5 6This lesson builds on tokens and pricing from `primer.ml.inference`, on the 7context window from `primer.agents.context` and on the release gate from 8`primer.agents.evals`. 9 10## Level 1: The practitioner's guide 11 12**In one sentence.** Cost engineering for a model-backed system means paying 13less per *successful* task, not per call, by taking the levers that can't 14hurt quality first and checking the ones that can against an eval. 15 16**When you need it.** When the bill grows faster than the value, when a 17latency budget is missed, or before the first big customer arrives and the 18price per task stops being a rounding error. The tell: you know your spend 19per month but not your cost per successful task, or you know cost per call 20but have never counted retries and human clean-up. The lesson's worked 21example shows why that number, and only that number, decides things: a 22small model at \$0.002 a call with a 60% success rate beats a large one at 23\$0.010 and 95% if failures can simply be retried (\$0.0033 against 24\$0.0105 per success), and loses by 7x if a person has to fix each failure 25(\$0.802 against \$0.110 per task, at an illustrative \$2 per fix). You don't 26need any of this for a prototype with ten users a day, and you shouldn't 27touch the levers that change behaviour (routing, semantic caching) until 28you have an eval to watch them with. All prices in this lesson are 29illustrative constants chosen for round arithmetic (a large model at \$5 30per million input tokens and \$25 per million output, a small one at \$1 31and \$5); check a provider's current price list for real numbers. 32 33**Your options.** Eight levers, from the ones that can't hurt quality to the 34ones that can: 35 36| Option | What it does | What it guarantees | What it costs | Where it lives | 37|---|---|---|---|---| 38| Prompt caching | Puts the stable part of the prompt first so the provider reuses its work | No behaviour change; the worked 10,000-token call falls from \$0.0625 to \$0.0265 with 8,000 tokens cached | A small write premium the first time; reordering the prompt | The provider's API | 39| Trim tokens | Compress tool results, send fewer retrieved chunks, ask for concise output | No change if you only cut what the next step never read; output is the expensive direction | Engineering, and care not to cut what was needed | Your prompt and code | 40| Parallel tool calls | Run independent tool calls at the same time | Wall-clock time falls to the slowest call (300 ms to 100 ms for three lookups); tokens unchanged | Concurrency in the agent loop | Your agent loop | 41| Exact response cache | The same question, ignoring case and spaces, returns the stored answer | Never wrong until the facts change | A time-to-live to manage | Your code | 42| Batch API | Non-interactive work submitted in bulk at about half price | Half price, results within hours | Waiting; only for work nobody is waiting on | The provider's API | 43| Budgets and alerts | Caps steps and tokens per task and spend per tenant; alerts on a sudden jump | A confused agent stops instead of looping all night | Choosing the limits; a false alarm now and then | Your code | 44| Routing | Sends easy tasks (classify, extract, format) to a small model and hard ones to a large one | 53% saved on this lesson's ten-task workload, if the router is right | The first lever that can lose quality; needs an eval before and after | Your code, in front of every request | 45| Semantic cache | Answers a *similar* question from memory, using embeddings to judge similarity | The cheapest possible hit | Confidently wrong answers: on this lesson's pairs, 40% of different-intent questions are served from cache at a 0.95 threshold | Your code, plus an embedding model | 46 47A ninth, for later: train or distil a small model on the large model's 48answers so it can take more of the routed traffic (Hinton et al., 2015). 49 50**How to choose.** Measure first, then go down the table in order. 51 52- No measurement yet: instrument cost per task by step, by model, input 53 against output and cached against uncached, and put an eval set in place 54 to hold quality fixed. 55- A long, stable system prompt or tool list: prompt caching, today. It is 56 the only lever that halves a bill without changing a single answer. 57- Big tool results or many retrieved chunks: trim. Compress to the fields 58 the next step needs and send the top few reranked chunks. 59- Several independent lookups per turn: run them concurrently, and stream 60 the answer so users see progress. 61- Anything nobody is waiting on (nightly evals, backfills, bulk 62 classification): batch it. 63- Most traffic is simple: route, with the eval watching. A router that sends 64 hard tasks to the small model saves money and quietly loses quality. 65- Narrow, curated, FAQ-style traffic and nothing else: a semantic cache 66 with guards (identifiers, numbers and negation must match), a 67 time-to-live and a measured wrong-hit rate. On everyday questions no 68 threshold makes it safe. 69- Whatever you pick, judge it by cost per successful task, with retries and 70 human clean-up counted in. 71 72**What it costs.** The levers stack. On this lesson's ten-task workload the 73baseline is \$0.671; caching takes it to \$0.311 (2.2x), trimming tool output 74to \$0.186 (3.6x), concise output to \$0.171 (3.9x), routing to \$0.086 (7.8x) 75and batching the non-interactive share to \$0.080 (8.4x). The first three 76are free wins that don't touch quality; routing is the big one and the 77first that can. Caching is not free on the first request: Anthropic's 78prompt caching documentation, for one, prices a cache write at 1.25 times 79the base input price for a five-minute cache (2 times for an hour), reads at 80a tenth of it or less, and only caches prompts above a per-model minimum of 81a few hundred to a few thousand tokens. Batching costs time: the same 82provider quotes a 50% discount with most batches finishing within an hour. 83Parallel calls and streaming cost no tokens at all; they buy time and 84perceived speed. Budgets cost the occasional false alarm: this lesson's 85anomaly rule flags a task using more than the mean plus three standard 86deviations of recent usage, so a 10,000-token task against a history around 871,000 trips it at once. 88 89**What breaks.** 90 91- **Cost per call replacing cost per success.** The cheap model looks 92 cheaper until failures are priced. Count retries and fixes. 93- **Routing that loses quality quietly.** Nothing errors; the answers just 94 get worse. Gate the router with the eval, and re-check when the small 95 model changes. 96- **The semantic cache that answers the wrong question.** "How many sick 97 days do I get?" scores 0.97 against the vacation question and gets the 98 vacation answer. Near-misses score higher than real paraphrases, so no 99 threshold separates them. 100- **Stale cache hits.** The policy changed; the cache didn't. Give every 101 entry a time-to-live. 102- **A cache that never hits.** Something volatile (a timestamp, the user's 103 name) sits at the top of the prompt, so the prefix differs every call. 104 Stable content first. 105- **Trimming what was needed.** A tool result cut to 500 tokens loses the 106 field the next step reads. Compress by field, not by length alone. 107- **The runaway loop.** An agent calls the same tool with the same 108 arguments all night. Cap steps and tokens per task; alert on tokens per 109 task jumping. 110- **One tenant's surprise invoice.** A shared platform with no per-customer 111 cap turns one customer's bug into everyone's bill. 112 113**In the wild.** Prompt caching and batch processing are standard features 114of hosted APIs; Anthropic's documentation for both is linked in Further 115reading and gives the multipliers quoted above. FrugalGPT (Chen, Zaharia and 116Zou, 2023) named the three families, prompt adaptation, model approximation 117and cascades that try a cheap model first and escalate, and reported 118matching the best single model with up to 98% less cost on its benchmarks. 119Distillation (Hinton, Vinyals and Dean, 2015) is how the small model in the 120routing table gets good enough to take more traffic. Tooling: LiteLLM is an 121open-source gateway that puts one interface in front of many providers and 122adds routing with fallbacks, per-key and per-team budgets, spend tracking 123and caching; GPTCache is an open-source semantic cache built on embeddings 124and a vector store with a pluggable similarity evaluator, the design this 125lesson's `SemanticCache` reproduces in miniature, guards and all. Parallel 126tool use is part of the tool-calling protocol of the major APIs: the model 127asks for several tools in one turn and you return all the results in one 128message. 129 130**Go deeper.** Level 2 prices a request symbol by symbol, builds the router, 131both caches and the parallel loop in plain Python, derives the anomaly 132alert from a mean and a standard deviation, works the unit economics with 133every number, and stacks the levers to draw the 8.4x figure. If you only 134needed to know which lever to pull first, you are done. 135 136## Level 2: How it works, from scratch 137 138A restaurant doesn't cut costs by buying worse ingredients for every dish. 139It sends simple orders to the line cook and complex ones to the head chef, 140pre-chops what every dish shares, doesn't plate food nobody eats, cooks 141things in parallel, does bulk prep overnight when it's cheaper, and watches 142the one kitchen station whose bills suddenly spike. 143 144Every lever in this lesson is one of those moves. The measure that matters 145is **cost per successful task**, not cost per call: a cheap attempt that 146fails still has to be paid for, and so does fixing it. 147 148**All prices in this module are illustrative constants** chosen for round 149arithmetic (a "large" model at \$5 per million input tokens and \$25 per 150million output tokens; a "small" one at \$1 and \$5). They show the *shape* 151of the trade-offs; check your provider's current price list for real 152numbers. 153 154## How a request is priced 155 156**Everyday picture.** A taxi that charges one rate for the distance to your 157pickup and a higher rate for the ride itself. Models charge per **token** 158(a word piece, about 4 characters of English): one rate for tokens you send 159(**input**) and a higher rate for tokens the model writes (**output**), 160because writing happens one token at a time (see `primer.ml.inference`). 161 162**Worked example.** A large-model call with 10,000 input tokens and 500 163output tokens: 10,000 × \$5 / 1,000,000 = \$0.05 for input, plus 164500 × \$25 / 1,000,000 = \$0.0125 for output, total **\$0.0625**. If 8,000 of 165those input tokens are a stable prefix served from the **prompt cache** 166(reused work from an earlier identical beginning, billed here at 10% of the 167input price), input becomes 8,000 × \$0.50/M + 2,000 × \$5/M = \$0.014, and 168the call costs **\$0.0265**, 58% less. 169 170$$ 171\text{cost} = \frac{c \cdot \rho\, p_{\text{in}} + (n_{\text{in}} - c)\, p_{\text{in}} + n_{\text{out}}\, p_{\text{out}}}{10^6} 172$$ 173 174**Symbols** 175 176| Symbol | Meaning here | Worked example | 177|---|---|---| 178| $n_{\text{in}}$ | input tokens sent | 10,000 | 179| $c$ | input tokens served from the prompt cache | 8,000 | 180| $n_{\text{out}}$ | output tokens generated | 500 | 181| $p_{\text{in}}, p_{\text{out}}$ | price per million input / output tokens | \$5, \$25 | 182| $\rho$ | cached-read price as a fraction of normal input price (Greek *rho*) | 0.1 | 183| $10^6$ | prices are per million tokens | | 184 185**In words:** cached input at its discounted rate, plus the rest of the 186input at full rate, plus output at the output rate, all per million. 187 188**On the worked example:** (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ = 189(4,000 + 10,000 + 12,500) / 10⁶ = \$0.0265. 190 191**In Python:** 192 193```python 194n_in, c, n_out = 10_000, 8_000, 500 195p_in, p_out, rho = 5, 25, 0.1 196cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6 197round(cost, 4) # → 0.0265 198``` 199 200Two facts fall out: output tokens cost several times more than input (ask 201for concise answers), and a cached prefix is nearly free (put stable content 202first; see `primer.agents.context`). Caches usually charge a small premium 203the first time a prefix is written; this module ignores it for simplicity. 204 205**In code:** `request_cost` is the formula above, reading each model's 206`Price` from `PRICES`. 207 208## Route each task to the cheapest model that can do it 209 210**Everyday picture.** A hospital triage nurse: sprained ankles go to the 211nurse practitioner, chest pains to the cardiologist. Nobody sends every 212patient to the most expensive specialist. 213 214**Worked example.** The sample workload is 10 tasks: 7 simple (classify a 215ticket, extract a date) and 3 complex (plan a migration). All on the large 216model it costs \$0.671. Routing the 7 simple ones to the small model, which 217is 5x cheaper per token, brings it to \$0.314: **53% saved**, with the hard 218tasks still on the strong model. 219 220```mermaid 221flowchart LR 222 T[Incoming task] --> R{Router<br/>rules or a small classifier} 223 R -->|classify, extract,<br/>format, short| S[Small, fast model] 224 R -->|plan, analyze,<br/>multi-step reasoning| L[Large model] 225 S --> O[Result] 226 L --> O 227``` 228 229**Reading it:** the router is cheap (a few rules here; often a small 230classifier) and sits in front of every request. The whole saving comes from 231the left branch: most traffic in real systems is simple, and simple work 232runs well on small models. Validate the router with evals 233(`primer.agents.evals`): a router that sends hard tasks to the small model 234saves money and quietly loses quality. 235 236**In code:** `route` is the rule-based router; `workload_cost` prices a list 237of `WorkItem` tasks with any levers switched on; `routing_savings` compares 238all-large against routed on `SAMPLE_WORKLOAD`. 239 240## Response caching and semantic caching 241 242**Everyday picture.** A receptionist who's been asked "what's the wifi 243password?" a hundred times just answers from memory. That's an **exact 244response cache**: the same question (ignoring case and spaces) returns the 245stored answer with no model call. A **semantic cache** goes further and 246answers *similar* questions from memory, using embeddings (vectors where 247closeness means similar meaning, see `primer.ml.embeddings`) to decide 248"similar". That's where the receptionist starts giving the vacation-policy 249answer to someone asking about sick leave. 250 251**Worked example.** With the toy embedder in `primer.common.embedder`: 252 253| Stored question | New question | Similarity | Same intent? | 254|---|---|---|---| 255| How do I reset my password? | I forgot my password, how do I recover it? | 0.88 | yes | 256| How many vacation days do I get? | How many sick days do I get? | **0.97** | **no** | 257| Request a new laptop | Request a new monitor | **0.96** | **no** | 258| What does ERR-4012 mean? | What does ERR-4013 mean? | 0.56 | no | 259 260The two near-misses score *higher* than a genuine paraphrase. No single 261threshold separates them. Two cheap **guards** help: identifiers and 262numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match 263("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which 264share a topic and differ in a single word. 265 266```mermaid 267flowchart TD 268 Q[New question] --> E{Exact match<br/>in response cache?} 269 E -->|yes| A1[Return stored answer] 270 E -->|no| V[Embed and find the<br/>nearest stored question] 271 V --> T{Similarity above<br/>threshold?} 272 T -->|no| M[Call the model] 273 T -->|yes| G{IDs, numbers and<br/>negation identical?<br/>entry not expired?} 274 G -->|no| M 275 G -->|yes| A2[Return stored answer] 276 M --> S[Store the new answer] 277``` 278 279**Reading it:** the exact cache is checked first because it's free and never 280wrong. The semantic path adds two gates: the similarity threshold, and the 281guards. An expired entry (older than its **TTL**, time to live) is treated 282as a miss, so answers about changing facts don't go stale forever. 283 284 285 286**Reading it:** the x-axis is the similarity threshold; the blue line is the 287share of genuine paraphrases answered from cache (good), and the red line is 288the share of different-intent questions answered from cache (a wrong answer 289served confidently). Lowering the threshold raises both, but look where the 290lines sit: from 0.4 upward the red line is above the blue one, so the cache 291serves more wrong answers than right ones (60% of different-intent questions 292against at most 57% of paraphrases). Even at 0.95 it still answers 40% of 293them wrongly and only 14% of paraphrases, because "sick" vs "vacation" 294scores 0.97. On everyday questions no threshold makes semantic caching safe; 295it's safe only for narrow, curated FAQ-style traffic, with guards, a TTL, 296and a measured wrong-hit rate. 297 298**In code:** `ResponseCache` is the exact cache. `SemanticCache` is the 299semantic path: `SemanticCache.lookup` skips expired entries and entries whose 300identifiers or negation differ, then returns the nearest survivor, and 301`SemanticCache.get` applies the threshold. `semantic_cache_sweep` draws the 302figure from the labelled pairs in `CACHE_PAIRS`. 303 304## Trim tokens 305 306**Everyday picture.** Don't photocopy the whole binder when the colleague 307needs one page. Compress tool results to the fields the next step needs 308(`primer.agents.context.compress_tool_output`), send the top few reranked 309chunks instead of dozens, remove repeated boilerplate from system prompts, 310and ask for concise output, because output is the expensive direction. 311 312**In code:** `workload_cost` models trimming as cutting each task's tool 313output to 500 tokens, and concise output as cutting long answers by a third. 314 315## Run independent tool calls at the same time 316 317**Everyday picture.** Boil the pasta while the sauce simmers. Cooking them 318one after the other takes the sum of the times; together, the time of the 319slowest. 320 321**Worked example.** Three independent lookups of 100 ms each: sequentially 322about 300 ms; concurrently about 100 ms. 323 324```mermaid 325sequenceDiagram 326 participant A as Agent 327 participant T1 as Weather API 328 participant T2 as Calendar API 329 participant T3 as CRM API 330 Note over A,T3: Sequential: about 300 ms 331 A->>T1: call 332 T1-->>A: result 333 A->>T2: call 334 T2-->>A: result 335 A->>T3: call 336 T3-->>A: result 337 Note over A,T3: Parallel: about 100 ms 338 par 339 A->>T1: call 340 and 341 A->>T2: call 342 and 343 A->>T3: call 344 end 345 T1-->>A: result 346 T2-->>A: result 347 T3-->>A: result 348``` 349 350**Reading it:** in the top half each call waits for the previous result; in 351the bottom half the three requests leave together and the agent waits only 352for the slowest. Models can ask for several tools in one turn (parallel tool 353use); run them concurrently and return all results in one message. Parallel 354calls cut wall-clock time, not tokens. 355 356 357 358**Reading it:** each bar is one 100 ms tool call on a shared time axis. The 359sequential calls stack end to end, 300 ms in all; the parallel ones start 360together, so the whole batch finishes in the time of one. Run the lesson to 361measure it with `asyncio`: the real timings land within a few milliseconds 362of these bars. 363 364**In code:** `run_sequential` awaits each call before starting the next; 365`run_parallel` starts them all with Python's asyncio gather and waits once. 366 367**Streaming** (showing tokens as they're generated) doesn't reduce total 368time either, but users see progress immediately, which changes how fast the 369system *feels*. 370 371## Batch what isn't interactive 372 373**Everyday picture.** Sending the laundry out to be done overnight at half 374price instead of waiting at the express counter. Providers offer **batch 375APIs**: submit many requests, get results within hours (often within 24), 376typically at about half the price. Use them for anything no one is waiting 377on: nightly evals, backfilling document processing, bulk classification. 378 379**In code:** `batch_cost` prices a list of requests at the batch discount. 380 381## Budgets and alerts 382 383**Everyday picture.** A prepaid card with a hard limit, and a bank that 384texts you when a purchase looks nothing like your usual spending. 385 386A **task budget** caps steps and tokens per task, so a confused agent stops 387instead of looping all night. A **tenant spend cap** (a tenant is one 388customer organisation on a shared platform) stops one customer's runaway 389usage from becoming a surprise invoice. An **anomaly alert** fires when a 390task uses far more tokens than usual: 391 392$$ 393\text{alert if } x > \mu + z\,\sigma 394$$ 395 396**Symbols** 397 398| Symbol | Meaning here | Worked example | 399|---|---|---| 400| $x$ | tokens used by the task just finished | 10,000 | 401| $\mu$ | mean tokens per task over this tenant's recent history (Greek *mu*) | 1,000 | 402| $\sigma$ | standard deviation of that history: the typical distance from the mean (Greek *sigma*) | ≈ 71 | 403| $z$ | how many standard deviations count as unusual | 3 | 404 405**In words:** alert when a task uses more tokens than the usual amount plus 406three times the usual spread. 407 408**On the worked example:** history 1000, 1100, 900, 1050, 950, 1000 has mean 4091,000 and standard deviation ≈ 71 (the squared distances from the mean add 410up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7), 411so the line is 1,000 + 3 × 70.7 ≈ 1,212. 412A 10,000-token task is far above it: alert. Such jumps usually mean a loop 413or a bad deploy. 414 415**In Python:** 416 417```python 418import statistics 419history = [1000, 1100, 900, 1050, 950, 1000] 420mu = statistics.mean(history) 421# divides by 6 - 1, as above 422sigma = statistics.stdev(history) 423mu, round(sigma, 1) # → (1000, 70.7) 424z = 3 425# the alert line 426round(mu + z * sigma) # → 1212 427# alert? 42810_000 > mu + z * sigma # → True 429``` 430 431**In code:** `TaskBudget.charge` counts each step's tokens and raises 432`BudgetExceeded` at either limit; `TenantSpend.record` adds a finished task's 433cost to its tenant's total and returns a cap alert or an anomaly alert (the 434formula above). 435 436## Unit economics: cost per successful task 437 438**Everyday picture.** A cheap printer that jams on 40% of pages isn't 439cheap once you count the wasted paper and the time spent clearing jams. 440 441$$ 442\text{cost per success (retry until it works)} = \frac{c}{p} 443\qquad 444\text{cost per task (a person fixes failures)} = c + (1 - p)\,h 445$$ 446 447**Symbols** 448 449| Symbol | Meaning here | Small model | Large model | 450|---|---|---|---| 451| $c$ | model cost of one attempt | \$0.002 | \$0.010 | 452| $p$ | chance an attempt succeeds | 0.60 | 0.95 | 453| $1/p$ | expected number of attempts until one succeeds | 1.67 | 1.05 | 454| $h$ | cost of a person fixing one failure (illustrative: a few minutes of staff time) | \$2.00 | \$2.00 | 455 456**In words:** if you can simply retry, you pay one attempt's cost for every 457expected attempt; if failures need a person, every task pays its attempt 458plus the chance of failure times the cost of the fix. 459 460**On the worked example:** retrying, the small model costs 0.002 / 0.60 = 461\$0.0033 per success and the large one 0.010 / 0.95 = \$0.0105, so the small 462model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 = 463**\$0.802** per task and the large one 0.010 + 0.05 × 2.00 = **\$0.110**: the 464"cheap" model is 7x more expensive. 465 466**In Python:** 467 468```python 469def per_success(c, p): 470 # retry until it works 471 return c / p 472def per_task(c, p, h=2.00): 473 # a person fixes each failure 474 return c + (1 - p) * h 475round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4) # → (0.0033, 0.0105) 476round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3) # → (0.802, 0.11) 477# small vs. large, with cleanup 478round(per_task(0.002, 0.60) / per_task(0.010, 0.95)) # → 7 479``` 480 481 482 483**Reading it:** the left pair of bars assumes failures can be retried 484automatically: the small model wins. The right pair assumes a person has to 485fix each failure: the large model wins by a mile. Which world you're in 486decides which model is cheaper, and the model's price per token barely 487matters in the second one. 488 489**In code:** `cost_per_success_with_retries` is $c/p$ and 490`cost_per_success_with_cleanup` is $c + (1-p)\,h$. 491 492## Putting it together: a 5x plan, in order 493 494Order the levers so the ones that can't hurt quality come first: 495 4961. **Prompt caching**: reorder the prompt so the stable prefix is reused. 497 No behaviour change. 4982. **Trim tokens**: compress tool results and send fewer chunks. 4993. **Concise output**: ask for shorter answers where length adds nothing. 5004. **Route easy tasks to the small model**: the first lever that *can* 501 change quality, so it's validated with evals before and after. 5025. **Batch** the non-interactive share. 503 504 505 506**Reading it:** each bar is the cost of the same 10-task workload after 507applying every lever up to and including that one; the label is the 508cumulative reduction. Caching alone roughly halves it, trimming and concise 509output take it past 3x, routing takes it past 7x, and batching adds the 510last few percent. The three levers after the baseline (caching, trimming, 511concise output) are the free wins that don't touch quality. 512 513**In code:** `five_x_plan` switches the levers on one at a time, in this 514order, and reports the cost and cumulative reduction after each. 515 516## In 20 seconds 517- Measure cost per *successful* task, not per call; include retries and 518 human cleanup. 519- Free wins first: prompt caching (stable prefix first), trimming tool 520 output and chunks, concise output, batch APIs for non-interactive work. 521- Then route easy tasks to small models, validated with evals. 522- Run independent tool calls in parallel; stream to improve perceived 523 speed. 524- Budgets per task, spend caps per tenant, and alerts on sudden jumps in 525 tokens per task. 526- Semantic caches can serve confidently wrong answers; use them only on 527 narrow traffic, with guards and a measured wrong-hit rate. 528 529## Self-test questions 530 531**How would you cut the cost per task by 5x without hurting quality?** 532First measure: cost per successful task, broken down by step, model, input 533vs. output, and cached vs. uncached tokens, with an eval set to hold 534quality fixed. Then, in order: restructure prompts for prompt caching; 535compress tool results and send only the top reranked chunks; ask for 536concise output; route simple steps (classification, extraction, 537formatting) to a small model, checking the eval before and after; move 538anything non-interactive to a batch API; cap steps and tokens per task. 539On the sample workload here that's about 8x, and every step after the 540first three is checked against the eval set. 541 542**Why can a cheaper model cost more?** 543Because failures cost money too: retries, human cleanup, lost customers. 544At 60% success with a \$2 human fix, a \$0.002 call costs \$0.80 per task; 545a \$0.010 call at 95% costs \$0.11. 546 547**What are the risks of a semantic cache?** 548Returning a confident, wrong answer to a question that's similar but not 549the same ("sick days" vs "vacation days"), and serving stale answers after 550facts change. Mitigate with high thresholds, exact-match guards on IDs, 551numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit 552rate on labelled pairs before enabling it. 553 554**Your average tokens per task doubled overnight. What do you check?** 555Whether a deploy changed a prompt or a tool (a bigger tool output, a new 556retrieval setting), whether an agent is looping (the same tool called with 557the same arguments), and whether the prompt cache hit rate dropped 558(something volatile moved to the top). Traces make this a lookup rather 559than a guess (`primer.agents.observability`). 560 561## The papers behind this lesson 562 563- Hinton, Vinyals & Dean, *Distilling the Knowledge in a Neural Network* 564 (2015), https://arxiv.org/abs/1503.02531. Showed how to train a small 565 model to imitate a large one, the technique behind making the cheap model 566 good enough to take more of the routed traffic. 567 [annotated companion](../../papers/distillation.html) 568- Chen, Zaharia & Zou, *FrugalGPT: How to Use Large Language Models While 569 Reducing Cost and Improving Performance* (2023), 570 https://arxiv.org/abs/2305.05176. Studied prompt adaptation, caching and 571 cascades that try cheap models first and escalate to expensive ones only 572 when needed. 573 574## Further reading 575- Anthropic prompt caching docs: https://docs.claude.com/en/docs/build-with-claude/prompt-caching 576- Anthropic Message Batches docs: https://docs.claude.com/en/docs/build-with-claude/batch-processing 577- Anthropic, parallel tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use 578- Python `asyncio.gather`: https://docs.python.org/3/library/asyncio-task.html#asyncio.gather 579""" 580 581from __future__ import annotations 582 583import asyncio 584import re 585import statistics 586import time 587from dataclasses import dataclass, field 588from typing import Any, Callable 589 590import numpy as np 591 592from primer._show import banner, say, table, takeaway 593from primer.common.embedder import ConceptEmbedder 594 595# --------------------------------------------------------------------------- 596# 1. Pricing (ILLUSTRATIVE constants, not any vendor's price list) 597# --------------------------------------------------------------------------- 598 599 600@dataclass(frozen=True) 601class Price: 602 input_per_m: float # dollars per million input tokens 603 output_per_m: float # dollars per million output tokens 604 605 606PRICES: dict[str, Price] = {"small": Price(1.0, 5.0), "large": Price(5.0, 25.0)} 607CACHE_READ_MULTIPLIER = 0.1 # cached input billed at 10% of the input price 608BATCH_MULTIPLIER = 0.5 # batch interface at half price 609 610 611def request_cost(model: str, input_tokens: int, output_tokens: int, cached_tokens: int = 0) -> float: 612 """Dollar cost of one call. `cached_tokens` of the input are billed at the cached rate.""" 613 p = PRICES[model] 614 uncached = input_tokens - cached_tokens 615 return (cached_tokens * CACHE_READ_MULTIPLIER * p.input_per_m + uncached * p.input_per_m + output_tokens * p.output_per_m) / 1e6 616 617 618def batch_cost(requests: list[tuple[str, int, int]]) -> float: 619 """Cost of (model, input, output) requests submitted through a batch interface.""" 620 return BATCH_MULTIPLIER * sum(request_cost(m, i, o) for m, i, o in requests) 621 622 623# --------------------------------------------------------------------------- 624# 2. Routing 625# --------------------------------------------------------------------------- 626 627# Words that suggest multi-step reasoning. A real router is often a small 628# classifier trained on labelled traffic; rules are the transparent start. 629HARD_SIGNALS = ("plan", "analy", "design", "strategy", "compare", "debug", "reason", "multi-step", "trade-off", "why") 630 631 632def route(task_text: str, max_simple_chars: int = 400) -> str: 633 """Return "small" or "large" for a task.""" 634 text = task_text.lower() 635 if len(text) > max_simple_chars or any(s in text for s in HARD_SIGNALS): 636 return "large" 637 return "small" 638 639 640@dataclass 641class WorkItem: 642 text: str 643 output_tokens: int 644 interactive: bool = True 645 # Input is split so the levers have something to act on. 646 stable_prefix: int = 8_000 # system prompt + tool definitions, identical every call 647 tool_output: int = 3_000 # verbose tool results that could be compressed 648 rest: int = 1_000 # the question and its specific context 649 650 651SAMPLE_WORKLOAD: list[WorkItem] = ( 652 [WorkItem("Classify this ticket as billing, technical or other", 150, interactive=False) for _ in range(4)] 653 + [WorkItem("Extract the invoice date and total from this email", 150) for _ in range(3)] 654 + [WorkItem("Plan the migration of our billing system and analyze the risks of each step", 600) for _ in range(3)] 655) 656 657 658def workload_cost(work: list[WorkItem], cache: bool = False, trim: bool = False, concise: bool = False, 659 routed: bool = False, batch: bool = False) -> float: 660 """Cost of a workload with any combination of levers switched on.""" 661 total = 0.0 662 for w in work: 663 tool = 500 if trim else w.tool_output # compress tool output to the needed fields 664 out = int(w.output_tokens * 2 / 3) if concise and w.output_tokens > 300 else w.output_tokens 665 n_in = w.stable_prefix + tool + w.rest 666 model = route(w.text) if routed else "large" 667 c = request_cost(model, n_in, out, cached_tokens=w.stable_prefix if cache else 0) 668 if batch and not w.interactive: 669 c *= BATCH_MULTIPLIER 670 total += c 671 return total 672 673 674def routing_savings(work: list[WorkItem] | None = None) -> dict[str, float]: 675 work = work or SAMPLE_WORKLOAD 676 all_large = workload_cost(work) 677 routed = workload_cost(work, routed=True) 678 return {"all_large": all_large, "routed": routed, "savings": 1 - routed / all_large} 679 680 681# --------------------------------------------------------------------------- 682# 3. Response cache and semantic cache 683# --------------------------------------------------------------------------- 684 685 686def _normalize(q: str) -> str: 687 return " ".join(q.lower().split()) 688 689 690class ResponseCache: 691 """Exact-match cache: free and never wrong, but misses every paraphrase.""" 692 693 def __init__(self) -> None: 694 self._store: dict[str, str] = {} 695 696 def put(self, question: str, answer: str) -> None: 697 self._store[_normalize(question)] = answer 698 699 def get(self, question: str) -> str | None: 700 return self._store.get(_normalize(question)) 701 702 703_NEGATION = re.compile(r"\b(not|no|never|don't|dont|doesn't|isn't|can't|won't|cannot)\b") 704 705 706def _identifiers(text: str) -> set[str]: 707 """Tokens containing a digit: order ids, error codes, amounts, dates.""" 708 return {t for t in re.findall(r"[a-z0-9]+(?:-[a-z0-9]+)*", text.lower()) if any(c.isdigit() for c in t)} 709 710 711def _negated(text: str) -> bool: 712 return bool(_NEGATION.search(text.lower())) 713 714 715@dataclass 716class _Entry: 717 vector: np.ndarray 718 question: str 719 answer: str 720 stored_at: float 721 722 723class SemanticCache: 724 """Answer similar questions from memory, with guards against the obvious wrong hits. 725 726 Guards: identifiers and numbers must match exactly, negation must match, 727 and entries expire after `ttl_seconds`. None of this can separate 728 "sick days" from "vacation days"; only narrow scoping and a measured 729 wrong-hit rate can. 730 """ 731 732 def __init__(self, embedder: Any = None, threshold: float = 0.9, ttl_seconds: float | None = None, 733 clock: Callable[[], float] = time.monotonic): 734 self.embedder = embedder or ConceptEmbedder() 735 self.threshold = threshold 736 self.ttl = ttl_seconds 737 self.clock = clock 738 self.entries: list[_Entry] = [] 739 740 def put(self, question: str, answer: str) -> None: 741 self.entries.append(_Entry(self.embedder.encode(question), question, answer, self.clock())) 742 743 def lookup(self, question: str) -> tuple[_Entry | None, float]: 744 """Nearest stored entry that passes the guards, and its similarity.""" 745 q = self.embedder.encode(question) 746 best, best_sim = None, -1.0 747 for e in self.entries: 748 if self.ttl is not None and self.clock() - e.stored_at > self.ttl: 749 continue 750 if _identifiers(e.question) != _identifiers(question) or _negated(e.question) != _negated(question): 751 continue 752 sim = float(e.vector @ q) # vectors are unit length, so dot product = cosine similarity 753 if sim > best_sim: 754 best, best_sim = e, sim 755 return best, best_sim 756 757 def get(self, question: str) -> str | None: 758 entry, sim = self.lookup(question) 759 return entry.answer if entry is not None and sim >= self.threshold else None 760 761 762# (stored question, new question, same intent?) 763CACHE_PAIRS: list[tuple[str, str, bool]] = [ 764 ("How do I reset my password?", "I forgot my password, how do I recover it?", True), 765 ("How do I reset my password?", "How do I reset my VPN password?", False), 766 ("How many vacation days do I get?", "How much PTO do I get per year?", True), 767 ("How many vacation days do I get?", "How many sick days do I get?", False), 768 ("How do I get reimbursed for car mileage?", "automobile expense reimbursement", True), 769 ("Cancel my order", "Don't cancel my order", False), 770 ("What does ERR-4012 mean?", "What does ERR-4013 mean?", False), 771 ("How do I set up the VPN?", "VPN setup steps", True), 772 ("Report a phishing email", "How do I report a suspicious email?", True), 773 ("Request a new laptop", "Request a new monitor", False), 774 ("What is the travel per-diem?", "How much is the daily meal per-diem when traveling?", True), 775 ("Printer not working", "printing fails on floor printer", True), 776] 777 778 779def semantic_cache_sweep(thresholds: list[float]) -> list[dict[str, float]]: 780 """For each threshold: share of paraphrases served from cache, share of near-misses served wrongly.""" 781 out = [] 782 for th in thresholds: 783 useful = wrong = 0 784 for stored, new, same in CACHE_PAIRS: 785 cache = SemanticCache(threshold=th) 786 cache.put(stored, "cached answer") 787 hit = cache.get(new) is not None 788 useful += hit and same 789 wrong += hit and not same 790 n_same = sum(1 for p in CACHE_PAIRS if p[2]) 791 out.append({"threshold": th, "useful_hit_rate": useful / n_same, "wrong_hit_rate": wrong / (len(CACHE_PAIRS) - n_same)}) 792 return out 793 794 795# --------------------------------------------------------------------------- 796# 4. Parallel tool calls 797# --------------------------------------------------------------------------- 798 799 800async def _fake_tool(i: int, delay: float, t0: float, log: list) -> str: 801 start = time.perf_counter() - t0 802 await asyncio.sleep(delay) # stands in for a network call 803 log.append((i, start, time.perf_counter() - t0)) 804 return f"tool-{i}" 805 806 807async def run_sequential(delays: list[float]) -> dict[str, Any]: 808 t0, log = time.perf_counter(), [] 809 results = [await _fake_tool(i, d, t0, log) for i, d in enumerate(delays)] 810 return {"seconds": time.perf_counter() - t0, "results": results, "spans": sorted(log)} 811 812 813async def run_parallel(delays: list[float]) -> dict[str, Any]: 814 t0, log = time.perf_counter(), [] 815 # gather starts every call before awaiting any, and returns results in call order. 816 results = await asyncio.gather(*(_fake_tool(i, d, t0, log) for i, d in enumerate(delays))) 817 return {"seconds": time.perf_counter() - t0, "results": list(results), "spans": sorted(log)} 818 819 820# --------------------------------------------------------------------------- 821# 5. Budgets and spend alerts 822# --------------------------------------------------------------------------- 823 824 825class BudgetExceeded(RuntimeError): 826 pass 827 828 829@dataclass 830class TaskBudget: 831 """Hard per-task limits. The agent loop calls charge() before every model call.""" 832 833 max_steps: int 834 max_tokens: int 835 steps: int = 0 836 tokens: int = 0 837 838 def charge(self, tokens: int) -> None: 839 if self.steps + 1 > self.max_steps: 840 raise BudgetExceeded(f"step limit {self.max_steps} reached") 841 if self.tokens + tokens > self.max_tokens: 842 raise BudgetExceeded(f"token limit {self.max_tokens} would be exceeded ({self.tokens} + {tokens})") 843 self.steps += 1 844 self.tokens += tokens 845 846 847@dataclass 848class TenantSpend: 849 """Per-tenant monthly spend cap plus an anomaly alert on tokens per task.""" 850 851 monthly_cap: float 852 z: float = 3.0 853 min_history: int = 5 854 spent: dict[str, float] = field(default_factory=dict) 855 history: dict[str, list[int]] = field(default_factory=dict) 856 857 def record(self, tenant: str, cost: float, tokens: int) -> dict[str, str]: 858 """Record one finished task; return alerts keyed by kind ("cap", "anomaly").""" 859 alerts: dict[str, str] = {} 860 self.spent[tenant] = self.spent.get(tenant, 0.0) + cost 861 if self.spent[tenant] > self.monthly_cap: 862 alerts["cap"] = f"{tenant} spent {self.spent[tenant]:.2f} of a {self.monthly_cap:.2f} cap" 863 hist = self.history.setdefault(tenant, []) 864 if len(hist) >= self.min_history: 865 mu, sigma = statistics.mean(hist), statistics.stdev(hist) 866 if tokens > mu + self.z * max(sigma, 1.0): 867 alerts["anomaly"] = f"{tokens} tokens vs usual {mu:.0f} ± {sigma:.0f}" 868 hist.append(tokens) 869 return alerts 870 871 872# --------------------------------------------------------------------------- 873# 6. Unit economics and the plan 874# --------------------------------------------------------------------------- 875 876 877def cost_per_success_with_retries(cost_per_attempt: float, success_rate: float) -> float: 878 """Retry until success: expected attempts are 1/p.""" 879 return cost_per_attempt / success_rate 880 881 882def cost_per_success_with_cleanup(cost_per_attempt: float, success_rate: float, human_fix_cost: float) -> float: 883 """One attempt per task; a person fixes each failure.""" 884 return cost_per_attempt + (1 - success_rate) * human_fix_cost 885 886 887def five_x_plan(work: list[WorkItem] | None = None) -> list[dict[str, Any]]: 888 """Apply the levers cumulatively, free ones first. Returns cost after each.""" 889 work = work or SAMPLE_WORKLOAD 890 levers = [ 891 ("baseline: everything on the large model", {}), 892 ("prompt caching", {"cache": True}), 893 ("trim tool output", {"cache": True, "trim": True}), 894 ("concise output", {"cache": True, "trim": True, "concise": True}), 895 ("route easy tasks to the small model", {"cache": True, "trim": True, "concise": True, "routed": True}), 896 ("batch non-interactive work", {"cache": True, "trim": True, "concise": True, "routed": True, "batch": True}), 897 ] 898 base = workload_cost(work) 899 steps = [] 900 for name, kw in levers: 901 c = workload_cost(work, **kw) 902 steps.append({"lever": name, "cost": c, "factor": base / c}) 903 return steps 904 905 906# --------------------------------------------------------------------------- 907# Figures and demo 908# --------------------------------------------------------------------------- 909 910 911def figures() -> dict[str, Any]: 912 import matplotlib 913 914 matplotlib.use("Agg") 915 import matplotlib.pyplot as plt 916 917 figs: dict[str, Any] = {} 918 919 # 1. Unit economics. 920 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 921 labels = ["retry until success", "person fixes failures"] 922 small = [cost_per_success_with_retries(0.002, 0.60), cost_per_success_with_cleanup(0.002, 0.60, 2.0)] 923 large = [cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)] 924 x = np.arange(2) 925 ax.bar(x - 0.2, small, 0.4, label="small: $0.002/attempt, 60% success", color="#c44e52") 926 ax.bar(x + 0.2, large, 0.4, label="large: $0.010/attempt, 95% success", color="#4c72b0") 927 for xi, (s, lg) in enumerate(zip(small, large)): 928 ax.text(xi - 0.2, s, f"${s:.3f}", ha="center", va="bottom", fontsize=8) 929 ax.text(xi + 0.2, lg, f"${lg:.3f}", ha="center", va="bottom", fontsize=8) 930 ax.set_yscale("log") 931 ax.set_xticks(x, labels) 932 ax.set_ylabel("dollars per successful task (log scale)") 933 ax.set_title("The cheap model is only cheap if failures are free") 934 ax.legend(fontsize=8, loc="upper left") 935 fig.tight_layout() 936 figs["unit_economics"] = fig 937 938 # 2. The 5x plan. 939 steps = five_x_plan() 940 fig, ax = plt.subplots(figsize=(7.5, 3.8)) 941 ax.bar(range(len(steps)), [s["cost"] for s in steps], color=["#8c8c8c"] + ["#4c72b0"] * 3 + ["#dd8452"] * 2) 942 for i, s in enumerate(steps): 943 ax.text(i, s["cost"], f"{s['factor']:.1f}x", ha="center", va="bottom", fontsize=9) 944 ax.set_xticks(range(len(steps)), ["baseline", "+ caching", "+ trim", "+ concise", "+ routing", "+ batch"]) 945 ax.set_ylabel("workload cost ($)") 946 ax.set_title("Levers applied in order (blue: can't change quality; orange: validate with evals)") 947 fig.tight_layout() 948 figs["five_x"] = fig 949 950 # 3. Semantic cache threshold sweep. 951 ths = list(np.round(np.arange(0.3, 1.001, 0.025), 3)) 952 sweep = semantic_cache_sweep(ths) 953 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 954 ax.plot(ths, [s["useful_hit_rate"] for s in sweep], color="#4c72b0", label="paraphrases served from cache (good)") 955 ax.plot(ths, [s["wrong_hit_rate"] for s in sweep], color="#c44e52", label="different questions served from cache (wrong)") 956 ax.set_xlabel("similarity threshold") 957 ax.set_ylabel("share of pairs") 958 ax.set_title("Semantic cache: no threshold separates them cleanly") 959 ax.legend(fontsize=8) 960 fig.tight_layout() 961 figs["semantic_cache"] = fig 962 963 # 4. Parallel vs sequential timeline: the schedule each approach follows. Drawn from the delays, not 964 # a stopwatch, so the site doesn't change with a millisecond of jitter; the demo measures the real thing. 965 delays_ms = [100, 100, 100] 966 sequential = [(i, sum(delays_ms[:i]), sum(delays_ms[: i + 1])) for i in range(len(delays_ms))] 967 parallel = [(i, 0, d) for i, d in enumerate(delays_ms)] 968 fig, ax = plt.subplots(figsize=(6.5, 3)) 969 for row, (name, spans, color) in enumerate([("sequential", sequential, "#c44e52"), ("parallel", parallel, "#4c72b0")]): 970 for i, start, end in spans: 971 ax.barh(row * 4 + i, end - start, left=start, color=color) 972 ax.text(start + 2, row * 4 + i, f"tool {i}", va="center", fontsize=8, color="white") 973 ax.set_yticks([1, 5], ["sequential", "parallel"]) 974 ax.invert_yaxis() 975 ax.set_xlabel("milliseconds since the agent started the calls") 976 ax.set_title("Three independent 100 ms tool calls") 977 fig.tight_layout() 978 figs["parallel"] = fig 979 return figs 980 981 982def demo() -> None: 983 banner("1. How a call is priced (illustrative prices)") 984 table(["call", "cost"], [ 985 ("large, 10k in / 500 out", f"${request_cost('large', 10_000, 500):.4f}"), 986 ("same, 8k of input cached", f"${request_cost('large', 10_000, 500, cached_tokens=8_000):.4f}"), 987 ("same, via batch interface", f"${batch_cost([('large', 10_000, 500)]):.4f}"), 988 ("small, 10k in / 500 out", f"${request_cost('small', 10_000, 500):.4f}"), 989 ]) 990 991 banner("2. Routing") 992 r = routing_savings() 993 print(f"all large ${r['all_large']:.3f} -> routed ${r['routed']:.3f}: {r['savings']:.0%} saved") 994 print() 995 996 banner("3. Semantic caching: useful and dangerous") 997 cache = SemanticCache(threshold=0.85) 998 cache.put("How do I reset my password?", "Use the self-service portal.") 999 cache.put("How many vacation days do I get?", "20 days of PTO per year.") 1000 for q in ["I forgot my password, how do I recover it?", "How many sick days do I get?", "How do I reset my VPN password?"]: 1001 entry, sim = cache.lookup(q) 1002 print(f"{q!r:48} sim {sim:.2f} -> {cache.get(q)!r}") 1003 print() 1004 say("The sick-days question gets the vacation answer. That's the risk: similar is not the same.") 1005 1006 banner("4. Parallel tool calls") 1007 seq, par = asyncio.run(run_sequential([0.1] * 3)), asyncio.run(run_parallel([0.1] * 3)) 1008 print(f"sequential {seq['seconds'] * 1000:.0f} ms, parallel {par['seconds'] * 1000:.0f} ms") 1009 print() 1010 1011 banner("5. Budgets and alerts") 1012 spend = TenantSpend(monthly_cap=1000) 1013 for t in [1000, 1100, 900, 1050, 950, 1000, 10_000]: 1014 alerts = spend.record("acme", 0.01, t) 1015 if alerts: 1016 print(f"task with {t} tokens -> {alerts}") 1017 print() 1018 1019 banner("6. Unit economics") 1020 table(["model", "per success (retry)", "per task (human fixes)"], [ 1021 ("small $0.002 @ 60%", cost_per_success_with_retries(0.002, 0.6), cost_per_success_with_cleanup(0.002, 0.6, 2.0)), 1022 ("large $0.010 @ 95%", cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)), 1023 ]) 1024 1025 banner("7. The plan, free levers first") 1026 table(["lever", "workload cost", "cumulative"], [(s["lever"], f"${s['cost']:.4f}", f"{s['factor']:.1f}x") for s in five_x_plan()]) 1027 takeaway("Measure cost per successful task, take the free wins first, then route with evals watching quality.") 1028 1029 1030if __name__ == "__main__": 1031 demo()
601@dataclass(frozen=True) 602class Price: 603 input_per_m: float # dollars per million input tokens 604 output_per_m: float # dollars per million output tokens
612def request_cost(model: str, input_tokens: int, output_tokens: int, cached_tokens: int = 0) -> float: 613 """Dollar cost of one call. `cached_tokens` of the input are billed at the cached rate.""" 614 p = PRICES[model] 615 uncached = input_tokens - cached_tokens 616 return (cached_tokens * CACHE_READ_MULTIPLIER * p.input_per_m + uncached * p.input_per_m + output_tokens * p.output_per_m) / 1e6
Dollar cost of one call. cached_tokens of the input are billed at the cached rate.
619def batch_cost(requests: list[tuple[str, int, int]]) -> float: 620 """Cost of (model, input, output) requests submitted through a batch interface.""" 621 return BATCH_MULTIPLIER * sum(request_cost(m, i, o) for m, i, o in requests)
Cost of (model, input, output) requests submitted through a batch interface.
633def route(task_text: str, max_simple_chars: int = 400) -> str: 634 """Return "small" or "large" for a task.""" 635 text = task_text.lower() 636 if len(text) > max_simple_chars or any(s in text for s in HARD_SIGNALS): 637 return "large" 638 return "small"
Return "small" or "large" for a task.
641@dataclass 642class WorkItem: 643 text: str 644 output_tokens: int 645 interactive: bool = True 646 # Input is split so the levers have something to act on. 647 stable_prefix: int = 8_000 # system prompt + tool definitions, identical every call 648 tool_output: int = 3_000 # verbose tool results that could be compressed 649 rest: int = 1_000 # the question and its specific context
659def workload_cost(work: list[WorkItem], cache: bool = False, trim: bool = False, concise: bool = False, 660 routed: bool = False, batch: bool = False) -> float: 661 """Cost of a workload with any combination of levers switched on.""" 662 total = 0.0 663 for w in work: 664 tool = 500 if trim else w.tool_output # compress tool output to the needed fields 665 out = int(w.output_tokens * 2 / 3) if concise and w.output_tokens > 300 else w.output_tokens 666 n_in = w.stable_prefix + tool + w.rest 667 model = route(w.text) if routed else "large" 668 c = request_cost(model, n_in, out, cached_tokens=w.stable_prefix if cache else 0) 669 if batch and not w.interactive: 670 c *= BATCH_MULTIPLIER 671 total += c 672 return total
Cost of a workload with any combination of levers switched on.
691class ResponseCache: 692 """Exact-match cache: free and never wrong, but misses every paraphrase.""" 693 694 def __init__(self) -> None: 695 self._store: dict[str, str] = {} 696 697 def put(self, question: str, answer: str) -> None: 698 self._store[_normalize(question)] = answer 699 700 def get(self, question: str) -> str | None: 701 return self._store.get(_normalize(question))
Exact-match cache: free and never wrong, but misses every paraphrase.
724class SemanticCache: 725 """Answer similar questions from memory, with guards against the obvious wrong hits. 726 727 Guards: identifiers and numbers must match exactly, negation must match, 728 and entries expire after `ttl_seconds`. None of this can separate 729 "sick days" from "vacation days"; only narrow scoping and a measured 730 wrong-hit rate can. 731 """ 732 733 def __init__(self, embedder: Any = None, threshold: float = 0.9, ttl_seconds: float | None = None, 734 clock: Callable[[], float] = time.monotonic): 735 self.embedder = embedder or ConceptEmbedder() 736 self.threshold = threshold 737 self.ttl = ttl_seconds 738 self.clock = clock 739 self.entries: list[_Entry] = [] 740 741 def put(self, question: str, answer: str) -> None: 742 self.entries.append(_Entry(self.embedder.encode(question), question, answer, self.clock())) 743 744 def lookup(self, question: str) -> tuple[_Entry | None, float]: 745 """Nearest stored entry that passes the guards, and its similarity.""" 746 q = self.embedder.encode(question) 747 best, best_sim = None, -1.0 748 for e in self.entries: 749 if self.ttl is not None and self.clock() - e.stored_at > self.ttl: 750 continue 751 if _identifiers(e.question) != _identifiers(question) or _negated(e.question) != _negated(question): 752 continue 753 sim = float(e.vector @ q) # vectors are unit length, so dot product = cosine similarity 754 if sim > best_sim: 755 best, best_sim = e, sim 756 return best, best_sim 757 758 def get(self, question: str) -> str | None: 759 entry, sim = self.lookup(question) 760 return entry.answer if entry is not None and sim >= self.threshold else None
Answer similar questions from memory, with guards against the obvious wrong hits.
Guards: identifiers and numbers must match exactly, negation must match,
and entries expire after ttl_seconds. None of this can separate
"sick days" from "vacation days"; only narrow scoping and a measured
wrong-hit rate can.
733 def __init__(self, embedder: Any = None, threshold: float = 0.9, ttl_seconds: float | None = None, 734 clock: Callable[[], float] = time.monotonic): 735 self.embedder = embedder or ConceptEmbedder() 736 self.threshold = threshold 737 self.ttl = ttl_seconds 738 self.clock = clock 739 self.entries: list[_Entry] = []
744 def lookup(self, question: str) -> tuple[_Entry | None, float]: 745 """Nearest stored entry that passes the guards, and its similarity.""" 746 q = self.embedder.encode(question) 747 best, best_sim = None, -1.0 748 for e in self.entries: 749 if self.ttl is not None and self.clock() - e.stored_at > self.ttl: 750 continue 751 if _identifiers(e.question) != _identifiers(question) or _negated(e.question) != _negated(question): 752 continue 753 sim = float(e.vector @ q) # vectors are unit length, so dot product = cosine similarity 754 if sim > best_sim: 755 best, best_sim = e, sim 756 return best, best_sim
Nearest stored entry that passes the guards, and its similarity.
780def semantic_cache_sweep(thresholds: list[float]) -> list[dict[str, float]]: 781 """For each threshold: share of paraphrases served from cache, share of near-misses served wrongly.""" 782 out = [] 783 for th in thresholds: 784 useful = wrong = 0 785 for stored, new, same in CACHE_PAIRS: 786 cache = SemanticCache(threshold=th) 787 cache.put(stored, "cached answer") 788 hit = cache.get(new) is not None 789 useful += hit and same 790 wrong += hit and not same 791 n_same = sum(1 for p in CACHE_PAIRS if p[2]) 792 out.append({"threshold": th, "useful_hit_rate": useful / n_same, "wrong_hit_rate": wrong / (len(CACHE_PAIRS) - n_same)}) 793 return out
For each threshold: share of paraphrases served from cache, share of near-misses served wrongly.
814async def run_parallel(delays: list[float]) -> dict[str, Any]: 815 t0, log = time.perf_counter(), [] 816 # gather starts every call before awaiting any, and returns results in call order. 817 results = await asyncio.gather(*(_fake_tool(i, d, t0, log) for i, d in enumerate(delays))) 818 return {"seconds": time.perf_counter() - t0, "results": list(results), "spans": sorted(log)}
Unspecified run-time error.
830@dataclass 831class TaskBudget: 832 """Hard per-task limits. The agent loop calls charge() before every model call.""" 833 834 max_steps: int 835 max_tokens: int 836 steps: int = 0 837 tokens: int = 0 838 839 def charge(self, tokens: int) -> None: 840 if self.steps + 1 > self.max_steps: 841 raise BudgetExceeded(f"step limit {self.max_steps} reached") 842 if self.tokens + tokens > self.max_tokens: 843 raise BudgetExceeded(f"token limit {self.max_tokens} would be exceeded ({self.tokens} + {tokens})") 844 self.steps += 1 845 self.tokens += tokens
Hard per-task limits. The agent loop calls charge() before every model call.
839 def charge(self, tokens: int) -> None: 840 if self.steps + 1 > self.max_steps: 841 raise BudgetExceeded(f"step limit {self.max_steps} reached") 842 if self.tokens + tokens > self.max_tokens: 843 raise BudgetExceeded(f"token limit {self.max_tokens} would be exceeded ({self.tokens} + {tokens})") 844 self.steps += 1 845 self.tokens += tokens
848@dataclass 849class TenantSpend: 850 """Per-tenant monthly spend cap plus an anomaly alert on tokens per task.""" 851 852 monthly_cap: float 853 z: float = 3.0 854 min_history: int = 5 855 spent: dict[str, float] = field(default_factory=dict) 856 history: dict[str, list[int]] = field(default_factory=dict) 857 858 def record(self, tenant: str, cost: float, tokens: int) -> dict[str, str]: 859 """Record one finished task; return alerts keyed by kind ("cap", "anomaly").""" 860 alerts: dict[str, str] = {} 861 self.spent[tenant] = self.spent.get(tenant, 0.0) + cost 862 if self.spent[tenant] > self.monthly_cap: 863 alerts["cap"] = f"{tenant} spent {self.spent[tenant]:.2f} of a {self.monthly_cap:.2f} cap" 864 hist = self.history.setdefault(tenant, []) 865 if len(hist) >= self.min_history: 866 mu, sigma = statistics.mean(hist), statistics.stdev(hist) 867 if tokens > mu + self.z * max(sigma, 1.0): 868 alerts["anomaly"] = f"{tokens} tokens vs usual {mu:.0f} ± {sigma:.0f}" 869 hist.append(tokens) 870 return alerts
Per-tenant monthly spend cap plus an anomaly alert on tokens per task.
858 def record(self, tenant: str, cost: float, tokens: int) -> dict[str, str]: 859 """Record one finished task; return alerts keyed by kind ("cap", "anomaly").""" 860 alerts: dict[str, str] = {} 861 self.spent[tenant] = self.spent.get(tenant, 0.0) + cost 862 if self.spent[tenant] > self.monthly_cap: 863 alerts["cap"] = f"{tenant} spent {self.spent[tenant]:.2f} of a {self.monthly_cap:.2f} cap" 864 hist = self.history.setdefault(tenant, []) 865 if len(hist) >= self.min_history: 866 mu, sigma = statistics.mean(hist), statistics.stdev(hist) 867 if tokens > mu + self.z * max(sigma, 1.0): 868 alerts["anomaly"] = f"{tokens} tokens vs usual {mu:.0f} ± {sigma:.0f}" 869 hist.append(tokens) 870 return alerts
Record one finished task; return alerts keyed by kind ("cap", "anomaly").
878def cost_per_success_with_retries(cost_per_attempt: float, success_rate: float) -> float: 879 """Retry until success: expected attempts are 1/p.""" 880 return cost_per_attempt / success_rate
Retry until success: expected attempts are 1/p.
883def cost_per_success_with_cleanup(cost_per_attempt: float, success_rate: float, human_fix_cost: float) -> float: 884 """One attempt per task; a person fixes each failure.""" 885 return cost_per_attempt + (1 - success_rate) * human_fix_cost
One attempt per task; a person fixes each failure.
888def five_x_plan(work: list[WorkItem] | None = None) -> list[dict[str, Any]]: 889 """Apply the levers cumulatively, free ones first. Returns cost after each.""" 890 work = work or SAMPLE_WORKLOAD 891 levers = [ 892 ("baseline: everything on the large model", {}), 893 ("prompt caching", {"cache": True}), 894 ("trim tool output", {"cache": True, "trim": True}), 895 ("concise output", {"cache": True, "trim": True, "concise": True}), 896 ("route easy tasks to the small model", {"cache": True, "trim": True, "concise": True, "routed": True}), 897 ("batch non-interactive work", {"cache": True, "trim": True, "concise": True, "routed": True, "batch": True}), 898 ] 899 base = workload_cost(work) 900 steps = [] 901 for name, kw in levers: 902 c = workload_cost(work, **kw) 903 steps.append({"lever": name, "cost": c, "factor": base / c}) 904 return steps
Apply the levers cumulatively, free ones first. Returns cost after each.
912def figures() -> dict[str, Any]: 913 import matplotlib 914 915 matplotlib.use("Agg") 916 import matplotlib.pyplot as plt 917 918 figs: dict[str, Any] = {} 919 920 # 1. Unit economics. 921 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 922 labels = ["retry until success", "person fixes failures"] 923 small = [cost_per_success_with_retries(0.002, 0.60), cost_per_success_with_cleanup(0.002, 0.60, 2.0)] 924 large = [cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)] 925 x = np.arange(2) 926 ax.bar(x - 0.2, small, 0.4, label="small: $0.002/attempt, 60% success", color="#c44e52") 927 ax.bar(x + 0.2, large, 0.4, label="large: $0.010/attempt, 95% success", color="#4c72b0") 928 for xi, (s, lg) in enumerate(zip(small, large)): 929 ax.text(xi - 0.2, s, f"${s:.3f}", ha="center", va="bottom", fontsize=8) 930 ax.text(xi + 0.2, lg, f"${lg:.3f}", ha="center", va="bottom", fontsize=8) 931 ax.set_yscale("log") 932 ax.set_xticks(x, labels) 933 ax.set_ylabel("dollars per successful task (log scale)") 934 ax.set_title("The cheap model is only cheap if failures are free") 935 ax.legend(fontsize=8, loc="upper left") 936 fig.tight_layout() 937 figs["unit_economics"] = fig 938 939 # 2. The 5x plan. 940 steps = five_x_plan() 941 fig, ax = plt.subplots(figsize=(7.5, 3.8)) 942 ax.bar(range(len(steps)), [s["cost"] for s in steps], color=["#8c8c8c"] + ["#4c72b0"] * 3 + ["#dd8452"] * 2) 943 for i, s in enumerate(steps): 944 ax.text(i, s["cost"], f"{s['factor']:.1f}x", ha="center", va="bottom", fontsize=9) 945 ax.set_xticks(range(len(steps)), ["baseline", "+ caching", "+ trim", "+ concise", "+ routing", "+ batch"]) 946 ax.set_ylabel("workload cost ($)") 947 ax.set_title("Levers applied in order (blue: can't change quality; orange: validate with evals)") 948 fig.tight_layout() 949 figs["five_x"] = fig 950 951 # 3. Semantic cache threshold sweep. 952 ths = list(np.round(np.arange(0.3, 1.001, 0.025), 3)) 953 sweep = semantic_cache_sweep(ths) 954 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 955 ax.plot(ths, [s["useful_hit_rate"] for s in sweep], color="#4c72b0", label="paraphrases served from cache (good)") 956 ax.plot(ths, [s["wrong_hit_rate"] for s in sweep], color="#c44e52", label="different questions served from cache (wrong)") 957 ax.set_xlabel("similarity threshold") 958 ax.set_ylabel("share of pairs") 959 ax.set_title("Semantic cache: no threshold separates them cleanly") 960 ax.legend(fontsize=8) 961 fig.tight_layout() 962 figs["semantic_cache"] = fig 963 964 # 4. Parallel vs sequential timeline: the schedule each approach follows. Drawn from the delays, not 965 # a stopwatch, so the site doesn't change with a millisecond of jitter; the demo measures the real thing. 966 delays_ms = [100, 100, 100] 967 sequential = [(i, sum(delays_ms[:i]), sum(delays_ms[: i + 1])) for i in range(len(delays_ms))] 968 parallel = [(i, 0, d) for i, d in enumerate(delays_ms)] 969 fig, ax = plt.subplots(figsize=(6.5, 3)) 970 for row, (name, spans, color) in enumerate([("sequential", sequential, "#c44e52"), ("parallel", parallel, "#4c72b0")]): 971 for i, start, end in spans: 972 ax.barh(row * 4 + i, end - start, left=start, color=color) 973 ax.text(start + 2, row * 4 + i, f"tool {i}", va="center", fontsize=8, color="white") 974 ax.set_yticks([1, 5], ["sequential", "parallel"]) 975 ax.invert_yaxis() 976 ax.set_xlabel("milliseconds since the agent started the calls") 977 ax.set_title("Three independent 100 ms tool calls") 978 fig.tight_layout() 979 figs["parallel"] = fig 980 return figs
983def demo() -> None: 984 banner("1. How a call is priced (illustrative prices)") 985 table(["call", "cost"], [ 986 ("large, 10k in / 500 out", f"${request_cost('large', 10_000, 500):.4f}"), 987 ("same, 8k of input cached", f"${request_cost('large', 10_000, 500, cached_tokens=8_000):.4f}"), 988 ("same, via batch interface", f"${batch_cost([('large', 10_000, 500)]):.4f}"), 989 ("small, 10k in / 500 out", f"${request_cost('small', 10_000, 500):.4f}"), 990 ]) 991 992 banner("2. Routing") 993 r = routing_savings() 994 print(f"all large ${r['all_large']:.3f} -> routed ${r['routed']:.3f}: {r['savings']:.0%} saved") 995 print() 996 997 banner("3. Semantic caching: useful and dangerous") 998 cache = SemanticCache(threshold=0.85) 999 cache.put("How do I reset my password?", "Use the self-service portal.") 1000 cache.put("How many vacation days do I get?", "20 days of PTO per year.") 1001 for q in ["I forgot my password, how do I recover it?", "How many sick days do I get?", "How do I reset my VPN password?"]: 1002 entry, sim = cache.lookup(q) 1003 print(f"{q!r:48} sim {sim:.2f} -> {cache.get(q)!r}") 1004 print() 1005 say("The sick-days question gets the vacation answer. That's the risk: similar is not the same.") 1006 1007 banner("4. Parallel tool calls") 1008 seq, par = asyncio.run(run_sequential([0.1] * 3)), asyncio.run(run_parallel([0.1] * 3)) 1009 print(f"sequential {seq['seconds'] * 1000:.0f} ms, parallel {par['seconds'] * 1000:.0f} ms") 1010 print() 1011 1012 banner("5. Budgets and alerts") 1013 spend = TenantSpend(monthly_cap=1000) 1014 for t in [1000, 1100, 900, 1050, 950, 1000, 10_000]: 1015 alerts = spend.record("acme", 0.01, t) 1016 if alerts: 1017 print(f"task with {t} tokens -> {alerts}") 1018 print() 1019 1020 banner("6. Unit economics") 1021 table(["model", "per success (retry)", "per task (human fixes)"], [ 1022 ("small $0.002 @ 60%", cost_per_success_with_retries(0.002, 0.6), cost_per_success_with_cleanup(0.002, 0.6, 2.0)), 1023 ("large $0.010 @ 95%", cost_per_success_with_retries(0.010, 0.95), cost_per_success_with_cleanup(0.010, 0.95, 2.0)), 1024 ]) 1025 1026 banner("7. The plan, free levers first") 1027 table(["lever", "workload cost", "cumulative"], [(s["lever"], f"${s['cost']:.4f}", f"{s['factor']:.1f}x") for s in five_x_plan()]) 1028 takeaway("Measure cost per successful task, take the free wins first, then route with evals watching quality.")