primer.agents.failures
Why the hard ones fail: a catalogue of agent failures, and the fix for each
Run: python -m primer.agents.failures
This lesson builds on the agent loop from primer.agents.agent_loop, on
traces from primer.agents.observability and on the golden set from
primer.agents.evals.
Level 1: The practitioner's guide
In one sentence. Most agent failures in production are system failures with known fixes (too many steps, bad retrieval, ambiguous tools, a context that rots, loops, no evals, injected instructions, messy data, brittle integrations, no adoption), and this lesson is the catalogue: symptom, fix, and where the fix is built.
When you need it. When the agent works in demos and fails in production, and nobody can say why. The tell is the gap between the two: demos are three steps long and real tasks are twenty. This lesson's arithmetic explains the gap on its own. If each step succeeds 95% of the time, a ten-step task succeeds 60% of the time and a twenty-step task 36%; at 90% per step, twenty steps succeed 12% of the time; at 99%, 82%. At 95% per step, the most steps you can chain and still succeed nine times in ten is two. Nothing about the model changed between the demo and production; the exponent did. You need this catalogue before the first incident, as a checklist, and after every incident, to name what happened. You don't need the integration patterns (retries, breakers) for a tool that calls nothing outside your process, and you don't need loop detection for a workflow with a fixed number of steps.
Your options. Six defences, from the ones that protect one request to the ones that protect the whole system:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Retries with capped exponential backoff and jitter | Wait 0.1 s, then 0.2 s, doubling to a cap, with a random spread so clients don't retry in waves | A transient blip (a timeout, a rate limit) doesn't fail the task | Latency on the retried request; a duplicate write unless the write is idempotent | Around each outside call |
| Circuit breaker | After a run of failures, calls fail instantly for a cool-down, then one trial call probes for recovery | An outage fails fast and stays contained instead of tying every request up in timeouts (in this lesson's 50-second outage the dead service is called 5 times and 21 requests fail fast) | A fallback path for when the breaker is open | Around each dependency |
| Loop detection and budgets | Stop when the same tool is called with the same arguments three times, or when a step or token budget runs out | A confused agent stops and hands off instead of burning money all night | A false stop now and then; a handoff path | The agent loop |
| Fewer steps, verification, checkpoints | Shorten the chain, check a step's result where it happens and retry only that step, save state to resume from the middle | Attacks the exponent directly: errors stop travelling down the chain | Design effort; a check after each key step costs a call | The task design |
| Contract tests | Automated checks that an outside API still accepts and returns what your tool expects | An upstream change is caught by a test, not a user | Tests to write and keep current | Your test suite |
| Evals and traces | A golden set in CI and a trace of every run | Regressions don't ship silently, and every failure can be replayed to its first bad step | The two lessons before this one | Your pipeline and your monitoring |
How to choose. Read the symptom first; the catalogue maps each one to a fix.
- Long tasks fail and short ones don't: compounding error. Cut steps, verify after the key ones, checkpoint, and put a person at the critical points.
- Timeouts and expired credentials: brittle integrations. Retries with backoff for blips, a breaker for outages, contract tests for changes. Never retry an invalid request (it fails the same way every time), and never retry a write that isn't idempotent.
- The same call repeated, cost spiking: loops. Detect the repeat and set hard budgets.
- Confident wrong answers: bad retrieval, or messy source data. Fix search and ingestion before touching the prompt.
- Wrong tool or bad arguments: ambiguous tools. Fewer tools, precise descriptions, validated inputs.
- The agent obeys text it found in a document: prompt injection. Boundaries around untrusted content and privilege separation.
- It works and nobody uses it: adoption. Build with the people it serves.
- Whatever the symptom, look at the trace for the first failing step before changing anything, and add the case to the golden set.
What it costs. Retries cost time on the slow path: the lesson's waits are 0.1 s and 0.2 s before the third attempt succeeds, capped at 10 s however many attempts there are, and with full jitter each client waits a random time up to that cap. A breaker costs almost nothing while closed and saves a great deal while open: every request that would have waited 30 s on a dead service fails at once instead. Loop detection is a comparison of the last few calls. Verification after a step costs a model call per check, which is why it goes after the key steps, not all of them. The expensive fix is the design one, shortening chains, and it is also the one with the largest payoff: raising per-step reliability from 95% to 99% takes a twenty-step task from 36% to 82%. The lesson's figures draw each of these curves so the trade can be read off, not guessed.
What breaks.
- Retrying the invalid. A bad request retried five times is five times the load and the same error. Retry only the exceptions that can clear on their own.
- Retrying a write without an idempotency key. "Create order" twice is two orders.
- Retries in waves. Every client that failed at the same moment retries at the same moment, and the recovering service goes down again. Jitter.
- No breaker. One dead dependency, and every request waits out a full timeout; the whole agent looks down.
- A breaker with no fallback. Failing fast is only useful if the agent can do something else: use a cache, tell the user, hand off.
- Loop detection that only counts. Three calls of the same tool with refining queries is progress, not a loop. Compare arguments, not names.
- Fixing the prompt for a system fault. A retrieval miss, a tool error or a parsing bug looks like a model mistake from the answer alone. The trace shows which it was.
In the wild. The circuit breaker was named by Michael Nygard in
Release It! and written up by Martin Fowler with the three states this
lesson builds (closed, open, half-open); the retry guidance, capped
exponential backoff with jitter, is the AWS Builders' Library article the
lesson links. Libraries package both: in
Python, tenacity gives a retry decorator with exponential and random
exponential waits, stop conditions by attempts or elapsed time, and retry
only on chosen exception types, the same policy as this lesson's
retry_with_backoff. Anthropic's Building effective agents draws the
same conclusions from the other direction: start with simple prompts and
evaluation, add multi-step agents only where simpler designs fall short,
add complexity only when it demonstrably improves outcomes, and have the
agent pause for human feedback at checkpoints; its emphasis on tool design
as an interface problem is the fix for the ambiguous-tools row. The seven
rows this lesson doesn't build point to the lessons that do.
Go deeper. Level 2 builds three of the fixes in plain code: the compounding-error formula and its inverse (how many steps a target allows), retries with capped backoff and full jitter that raise invalid requests at once, a circuit breaker as a three-state machine replayed through an outage, and a loop detector that compares arguments. If you only needed to name the failure and find its fix, you are done.
Level 2: How it works, from scratch
Pilots learn from a catalogue of accidents: each entry names what went wrong, how it showed up in the cockpit, and the checklist item that now prevents it. This lesson is that catalogue for agents: ten failures that keep recurring in production systems, each with its symptom, its fix, and a pointer to the lesson that builds the fix. Three of them get their fix built right here: compounding error (arithmetic), brittle integrations (retries with backoff, and circuit breakers) and runaway loops (loop detection).
| Failure | Symptom | Fix | Lesson |
|---|---|---|---|
| Compounding error | Long tasks fail far more often than short ones | Fewer steps, checks after key steps, checkpoints | primer.agents.planning |
| Bad retrieval | Confident, wrong answers | Hybrid search, reranking, retrieval evals | primer.agents.rag |
| Ambiguous tools | Wrong tool, or bad arguments | Precise descriptions, fewer tools, validation | primer.agents.tools |
| Context rot | Quality drops as a session grows | Summarize, trim, restart with a state handoff | primer.agents.context |
| Loops and runaway | The same call repeated; cost spikes | Step budgets, loop detection, stop conditions | primer.agents.agent_loop |
| No evals | Regressions ship silently | Golden sets in CI, online monitoring | primer.agents.evals |
| Prompt injection | The agent obeys instructions found in data | Untrusted-content boundaries, privilege separation | primer.agents.guardrails |
| Messy enterprise data | Garbled tables, missing permissions | Invest in parsing; permission-aware retrieval | primer.agents.rag |
| Brittle integrations | Timeouts, expired credentials, API changes | Retries with backoff, circuit breakers, contract tests | primer.agents.failures |
| No adoption | It works, and nobody uses it | Build with users, show sources, easy human handoff | primer.agents.deployment |
In code: FAILURES is this table as data, one Failure record per row.
Compounding error
Everyday picture. A relay race where each baton pass succeeds 95% of the time. One pass is nearly safe; ten passes in a row drop the baton more often than you'd think.
Worked example. If each step of an agent's task succeeds 95% of the time, independently:
| Steps | Chance every step succeeds |
|---|---|
| 1 | 0.95 |
| 2 | 0.95 × 0.95 = 0.9025 |
| 10 | 0.95¹⁰ ≈ 0.599 |
| 20 | 0.95²⁰ ≈ 0.358 |
Level 3: the formula and its symbols
$$ P(\text{task succeeds}) = p^{\,n} $$
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| $p$ | chance a single step succeeds | 0.95 |
| $n$ | number of steps that must all succeed | 10 |
| $p^{\,n}$ | $p$ multiplied by itself $n$ times | 0.599 |
In words: the chance that every step succeeds is the per-step chance multiplied together once per step.
On the worked example: 0.95 multiplied by itself 10 times is 0.599, so a ten-step task at 95% per step fails four times in ten.
Level 3: in Python
In Python:
p, n = 0.95, 10
# p multiplied by itself n times
round(p ** n, 3) # → 0.599
# fails about four times in ten
round(1 - p ** n, 1) # → 0.4
flowchart LR S1[Step 1<br/>95%] --> S2[Step 2<br/>95%] --> S3[...] --> S10[Step 10<br/>95%] --> D[Done:<br/>60% of the time] S2 -.->|verify, retry<br/>just this step| S2
Reading it: every arrow is another chance to drop the baton, so success shrinks with each step. The dotted loop is the remedy: check a step's result where it happens and retry only that step, so an error doesn't travel down the chain. Fewer steps, verification after key steps, checkpoints to resume from the middle, and human review at critical points all attack the same exponent.
Reading it: the x-axis is the number of steps; each curve is a per-step reliability. At 99% per step a 20-step task still succeeds 82% of the time; at 95% it's 36%; at 90%, 12%. Small gains in per-step reliability are worth far more than they look, and demos with three steps say little about tasks with twenty.
In code: chain_success is $p^{\,n}$; max_steps_for runs it backwards,
returning the most steps you can chain and still reach a target success rate.
Brittle integrations: retries with backoff
Everyday picture. Calling a busy phone line: you don't redial every second. You wait a little, then longer, then longer still, and if everyone redials on the same schedule the line stays jammed, so you add a random pause (jitter).
Worked example. A service times out twice and then answers. With a base
delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt
succeeds. With a cap of 10 s, the waits never exceed 10 s however many
attempts there are. A request that's invalid (a ValueError here) fails
the same way every time, so it's raised immediately: retrying only adds
load.
Level 3: the formula and its symbols
$$ d_k = \min\bigl(D_{\max},\; b \cdot 2^{k}\bigr), \qquad k = 0, 1, 2, \ldots $$
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| $d_k$ | wait before retry number $k+1$ | $d_0$ = 0.1 s, $d_1$ = 0.2 s |
| $k$ | how many retries have already happened | 0, 1 |
| $b$ | base delay | 0.1 s |
| $2^{k}$ | doubles with every retry | 1, 2, 4, … |
| $D_{\max}$ | the cap on any single wait | 10 s |
In words: each wait doubles the previous one, starting from the base delay, but never exceeds the cap.
On the worked example: $d_0$ = min(10, 0.1 × 1) = 0.1 s and $d_1$ = min(10, 0.1 × 2) = 0.2 s.
Level 3: in Python
In Python:
b, D_max = 0.1, 10
# d_0 and d_1
[min(D_max, b * 2 ** k) for k in range(2)] # → [0.1, 0.2]
sequenceDiagram participant A as Agent tool participant S as Upstream API A->>S: request S--xA: timeout Note over A: wait 0.1 s A->>S: retry 1 S--xA: timeout Note over A: wait 0.2 s A->>S: retry 2 S-->>A: 200 OK
Reading it: time runs downward. Each failure is followed by a longer pause, which gives an overloaded service room to recover instead of being hammered. Writes must be idempotent (safe to repeat: sending the same request twice has the same effect as once, usually via an idempotency key) before you retry them, or a retried "create order" makes two orders.
Reading it: the x-axis is the attempt number; the black line is the capped exponential delay. The dots are five clients using full jitter (each waits a random time between zero and the capped delay): instead of all retrying at the same moment, their retries spread out, which is what lets a recovering service actually recover.
In code: backoff_delays lists the capped waits $d_k$;
retry_with_backoff calls a function, retries only the exceptions in
RETRYABLE with those waits (optionally with full jitter), and raises
anything else at once.
Brittle integrations: circuit breakers
Everyday picture. The breaker in your home's fuse box. When a circuit keeps shorting, it trips and cuts the power, instead of letting the wire overheat. After a while you flip it back on to test; if it trips again, it stays off.
A circuit breaker wraps calls to a dependency. After several consecutive failures it opens: calls fail immediately without touching the struggling service (fail fast), so the agent can fall back (use a cache, tell the user, hand off to a person) instead of hanging on timeouts. After a cool-down it goes half-open and lets one trial call through: success closes it, failure opens it again.
Worked example. Threshold 3, cool-down 30 s. Three timeouts in a row:
open. A fourth call at 5 s fails instantly with CircuitOpen and the
service isn't called at all. At 31 s, one trial call goes through; it
succeeds, so the breaker closes.
stateDiagram-v2 [*] --> Closed Closed --> Closed: success (reset the failure count) Closed --> Open: 3 failures in a row Open --> Open: calls rejected instantly Open --> HalfOpen: cool-down elapsed HalfOpen --> Closed: trial call succeeds HalfOpen --> Open: trial call fails
Reading it: in the normal state (Closed) calls flow through and failures are counted. Open is the protective state: nothing reaches the dependency. Half-open is a single, cautious test. The breaker turns a slow, cascading failure (every request waiting 30 s on a dead service) into a fast, contained one.
Reading it: the shaded band is when the upstream service is down; each marker is one request. Before the outage, calls succeed (green). The first three failures (red) open the breaker; after that, requests fail fast (grey) without reaching the service, apart from one trial call per cool-down. Once the service is back, the next trial succeeds and traffic resumes. The service received a handful of calls during the outage instead of all of them.
In code: CircuitBreaker.call is the state diagram: it raises
CircuitOpen while open, lets one trial call through after the cool-down,
and closes on success. simulate_outage replays the outage in the figure.
Contract tests (automated checks that an external API still accepts and returns what your tool expects) catch the third kind of brittleness, upstream API changes, before users do.
Loops and runaway
Everyday picture. A satnav that keeps rerouting you around the same block. The fix isn't a better map; it's noticing that you've passed the same corner three times.
Worked example. search("vpn"), search("vpn"), search("vpn"): the
same tool with the same arguments three times in a row. Nothing new can
come back, so the agent is looping. search("vpn"), search("vpn error"),
search("ERR-4012") is the same tool refining its query: progress.
Combine loop detection with hard step and token budgets
(primer.agents.cost.TaskBudget) so a confused agent stops and hands off
instead of burning money.
In code: is_looping reports whether the last few calls were the same
tool with identical arguments.
In 20 seconds
- Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60% of the time. Shorten chains and verify at key steps.
- Retry transient failures with capped exponential backoff and jitter; never retry invalid requests; make writes idempotent first.
- Put circuit breakers around dependencies so outages fail fast instead of cascading.
- Detect loops (same call, same arguments) and enforce step and token budgets.
- Most failures are system failures (retrieval, tools, data, evals, adoption), not model failures, and each has a known fix.
Self-test questions
Your agent works in demos and fails half the time in production. Where do you look first? At the length of real tasks versus demo tasks (compounding error), and at traces for the first failing step: retrieval misses, wrong tools or arguments, tool errors from integrations, context growth. Then fix per-step reliability where it's lowest, add verification after key steps, and put the failing cases into the eval set.
Why add jitter to retries? Without it, every client that failed at the same moment retries at the same moment, producing synchronized waves of load that keep an overloaded service down. Random delays spread the retries out.
When should you not retry? When the error can't go away on its own: invalid input, permission denied, not found. And never retry a non-idempotent write without an idempotency key, or you'll duplicate its effect.
What's the difference between a retry and a circuit breaker? A retry handles a blip for one request. A breaker handles an outage across many requests: it stops sending traffic to a failing dependency, fails fast so callers can fall back, and probes for recovery.
The papers behind this lesson
No single founding paper; these are engineering patterns. The canonical write-ups are Michael Nygard's Release It! (which named the circuit breaker) and the AWS Builders' Library article on backoff and jitter, both linked below.
Further reading
- Martin Fowler, CircuitBreaker: https://martinfowler.com/bliki/CircuitBreaker.html
- AWS Builders' Library, Timeouts, retries, and backoff with jitter: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
1r""" 2# Why the hard ones fail: a catalogue of agent failures, and the fix for each 3 4Run: `python -m primer.agents.failures` 5 6This lesson builds on the agent loop from `primer.agents.agent_loop`, on 7traces from `primer.agents.observability` and on the golden set from 8`primer.agents.evals`. 9 10## Level 1: The practitioner's guide 11 12**In one sentence.** Most agent failures in production are system failures 13with known fixes (too many steps, bad retrieval, ambiguous tools, a context 14that rots, loops, no evals, injected instructions, messy data, brittle 15integrations, no adoption), and this lesson is the catalogue: symptom, fix, 16and where the fix is built. 17 18**When you need it.** When the agent works in demos and fails in 19production, and nobody can say why. The tell is the gap between the two: 20demos are three steps long and real tasks are twenty. This lesson's 21arithmetic explains the gap on its own. If each step succeeds 95% of the 22time, a ten-step task succeeds 60% of the time and a twenty-step task 36%; 23at 90% per step, twenty steps succeed 12% of the time; at 99%, 82%. At 95% 24per step, the most steps you can chain and still succeed nine times in ten 25is two. Nothing about the model changed between the demo and production; 26the exponent did. You need this catalogue before the first incident, as a 27checklist, and after every incident, to name what happened. You don't need 28the integration patterns (retries, breakers) for a tool that calls nothing 29outside your process, and you don't need loop detection for a workflow 30with a fixed number of steps. 31 32**Your options.** Six defences, from the ones that protect one request to 33the ones that protect the whole system: 34 35| Option | What it does | What it guarantees | What it costs | Where it lives | 36|---|---|---|---|---| 37| Retries with capped exponential backoff and jitter | Wait 0.1 s, then 0.2 s, doubling to a cap, with a random spread so clients don't retry in waves | A transient blip (a timeout, a rate limit) doesn't fail the task | Latency on the retried request; a duplicate write unless the write is idempotent | Around each outside call | 38| Circuit breaker | After a run of failures, calls fail instantly for a cool-down, then one trial call probes for recovery | An outage fails fast and stays contained instead of tying every request up in timeouts (in this lesson's 50-second outage the dead service is called 5 times and 21 requests fail fast) | A fallback path for when the breaker is open | Around each dependency | 39| Loop detection and budgets | Stop when the same tool is called with the same arguments three times, or when a step or token budget runs out | A confused agent stops and hands off instead of burning money all night | A false stop now and then; a handoff path | The agent loop | 40| Fewer steps, verification, checkpoints | Shorten the chain, check a step's result where it happens and retry only that step, save state to resume from the middle | Attacks the exponent directly: errors stop travelling down the chain | Design effort; a check after each key step costs a call | The task design | 41| Contract tests | Automated checks that an outside API still accepts and returns what your tool expects | An upstream change is caught by a test, not a user | Tests to write and keep current | Your test suite | 42| Evals and traces | A golden set in CI and a trace of every run | Regressions don't ship silently, and every failure can be replayed to its first bad step | The two lessons before this one | Your pipeline and your monitoring | 43 44**How to choose.** Read the symptom first; the catalogue maps each one to 45a fix. 46 47- Long tasks fail and short ones don't: compounding error. Cut steps, 48 verify after the key ones, checkpoint, and put a person at the critical 49 points. 50- Timeouts and expired credentials: brittle integrations. Retries with 51 backoff for blips, a breaker for outages, contract tests for changes. 52 Never retry an invalid request (it fails the same way every time), and 53 never retry a write that isn't idempotent. 54- The same call repeated, cost spiking: loops. Detect the repeat and set 55 hard budgets. 56- Confident wrong answers: bad retrieval, or messy source data. Fix search 57 and ingestion before touching the prompt. 58- Wrong tool or bad arguments: ambiguous tools. Fewer tools, precise 59 descriptions, validated inputs. 60- The agent obeys text it found in a document: prompt injection. Boundaries 61 around untrusted content and privilege separation. 62- It works and nobody uses it: adoption. Build with the people it serves. 63- Whatever the symptom, look at the trace for the first failing step before 64 changing anything, and add the case to the golden set. 65 66**What it costs.** Retries cost time on the slow path: the lesson's waits 67are 0.1 s and 0.2 s before the third attempt succeeds, capped at 10 s 68however many attempts there are, and with full jitter each client waits a 69random time up to that cap. A breaker costs almost nothing while closed 70and saves a great deal while open: every request that would have waited 7130 s on a dead service fails at once instead. Loop detection is 72a comparison of the last few calls. Verification after a step costs a 73model call per check, which is why it goes after the key steps, not all of 74them. The expensive fix is the design one, shortening chains, and it is 75also the one with the largest payoff: raising per-step reliability from 7695% to 99% takes a twenty-step task from 36% to 82%. The lesson's figures 77draw each of these curves so the trade can be read off, not guessed. 78 79**What breaks.** 80 81- **Retrying the invalid.** A bad request retried five times is five times 82 the load and the same error. Retry only the exceptions that can clear 83 on their own. 84- **Retrying a write without an idempotency key.** "Create order" twice is 85 two orders. 86- **Retries in waves.** Every client that failed at the same moment retries 87 at the same moment, and the recovering service goes down again. Jitter. 88- **No breaker.** One dead dependency, and every request waits out a full 89 timeout; the whole agent looks down. 90- **A breaker with no fallback.** Failing fast is only useful if the agent 91 can do something else: use a cache, tell the user, hand off. 92- **Loop detection that only counts.** Three calls of the same tool with 93 refining queries is progress, not a loop. Compare arguments, not names. 94- **Fixing the prompt for a system fault.** A retrieval miss, a tool error 95 or a parsing bug looks like a model mistake from the answer alone. The 96 trace shows which it was. 97 98**In the wild.** The circuit breaker was named by Michael Nygard in 99*Release It!* and written up by Martin Fowler with the three states this 100lesson builds (closed, open, half-open); the retry guidance, capped 101exponential backoff with jitter, is the AWS Builders' Library article the 102lesson links. Libraries package both: in 103Python, tenacity gives a retry decorator with exponential and random 104exponential waits, stop conditions by attempts or elapsed time, and retry 105only on chosen exception types, the same policy as this lesson's 106`retry_with_backoff`. Anthropic's *Building effective agents* draws the 107same conclusions from the other direction: start with simple prompts and 108evaluation, add multi-step agents only where simpler designs fall short, 109add complexity only when it demonstrably improves outcomes, and have the 110agent pause for human feedback at checkpoints; its emphasis on tool design 111as an interface problem is the fix for the ambiguous-tools row. The seven 112rows this lesson doesn't build point to the lessons that do. 113 114**Go deeper.** Level 2 builds three of the fixes in plain code: the 115compounding-error formula and its inverse (how many steps a target allows), 116retries with capped backoff and full jitter that raise invalid requests 117at once, a circuit breaker as a three-state machine replayed through an 118outage, and a loop detector that compares arguments. If you only needed to 119name the failure and find its fix, you are done. 120 121## Level 2: How it works, from scratch 122 123Pilots learn from a catalogue of accidents: each entry names what went 124wrong, how it showed up in the cockpit, and the checklist item that now 125prevents it. This lesson is that catalogue for agents: ten failures that 126keep recurring in production systems, each with its symptom, its fix, and a 127pointer to the lesson that builds the fix. Three of them get their fix 128built right here: compounding error (arithmetic), brittle integrations 129(retries with backoff, and circuit breakers) and runaway loops (loop 130detection). 131 132| Failure | Symptom | Fix | Lesson | 133|---|---|---|---| 134| Compounding error | Long tasks fail far more often than short ones | Fewer steps, checks after key steps, checkpoints | `primer.agents.planning` | 135| Bad retrieval | Confident, wrong answers | Hybrid search, reranking, retrieval evals | `primer.agents.rag` | 136| Ambiguous tools | Wrong tool, or bad arguments | Precise descriptions, fewer tools, validation | `primer.agents.tools` | 137| Context rot | Quality drops as a session grows | Summarize, trim, restart with a state handoff | `primer.agents.context` | 138| Loops and runaway | The same call repeated; cost spikes | Step budgets, loop detection, stop conditions | `primer.agents.agent_loop` | 139| No evals | Regressions ship silently | Golden sets in CI, online monitoring | `primer.agents.evals` | 140| Prompt injection | The agent obeys instructions found in data | Untrusted-content boundaries, privilege separation | `primer.agents.guardrails` | 141| Messy enterprise data | Garbled tables, missing permissions | Invest in parsing; permission-aware retrieval | `primer.agents.rag` | 142| Brittle integrations | Timeouts, expired credentials, API changes | Retries with backoff, circuit breakers, contract tests | `primer.agents.failures` | 143| No adoption | It works, and nobody uses it | Build with users, show sources, easy human handoff | `primer.agents.deployment` | 144 145**In code:** `FAILURES` is this table as data, one `Failure` record per row. 146 147## Compounding error 148 149**Everyday picture.** A relay race where each baton pass succeeds 95% of 150the time. One pass is nearly safe; ten passes in a row drop the baton more 151often than you'd think. 152 153**Worked example.** If each step of an agent's task succeeds 95% of the 154time, independently: 155 156| Steps | Chance every step succeeds | 157|---|---| 158| 1 | 0.95 | 159| 2 | 0.95 × 0.95 = 0.9025 | 160| 10 | 0.95¹⁰ ≈ **0.599** | 161| 20 | 0.95²⁰ ≈ **0.358** | 162 163$$ 164P(\text{task succeeds}) = p^{\,n} 165$$ 166 167**Symbols** 168 169| Symbol | Meaning here | Worked example | 170|---|---|---| 171| $p$ | chance a single step succeeds | 0.95 | 172| $n$ | number of steps that must all succeed | 10 | 173| $p^{\,n}$ | $p$ multiplied by itself $n$ times | 0.599 | 174 175**In words:** the chance that every step succeeds is the per-step chance 176multiplied together once per step. 177 178**On the worked example:** 0.95 multiplied by itself 10 times is 0.599, so a 179ten-step task at 95% per step fails four times in ten. 180 181**In Python:** 182 183```python 184p, n = 0.95, 10 185# p multiplied by itself n times 186round(p ** n, 3) # → 0.599 187# fails about four times in ten 188round(1 - p ** n, 1) # → 0.4 189``` 190 191```mermaid 192flowchart LR 193 S1[Step 1<br/>95%] --> S2[Step 2<br/>95%] --> S3[...] --> S10[Step 10<br/>95%] --> D[Done:<br/>60% of the time] 194 S2 -.->|verify, retry<br/>just this step| S2 195``` 196 197**Reading it:** every arrow is another chance to drop the baton, so success 198shrinks with each step. The dotted loop is the remedy: check a step's 199result where it happens and retry only that step, so an error doesn't 200travel down the chain. Fewer steps, verification after key steps, 201checkpoints to resume from the middle, and human review at critical points 202all attack the same exponent. 203 204 205 206**Reading it:** the x-axis is the number of steps; each curve is a per-step 207reliability. At 99% per step a 20-step task still succeeds 82% of the time; 208at 95% it's 36%; at 90%, 12%. Small gains in per-step reliability are worth 209far more than they look, and demos with three steps say little about tasks 210with twenty. 211 212**In code:** `chain_success` is $p^{\,n}$; `max_steps_for` runs it backwards, 213returning the most steps you can chain and still reach a target success rate. 214 215## Brittle integrations: retries with backoff 216 217**Everyday picture.** Calling a busy phone line: you don't redial every 218second. You wait a little, then longer, then longer still, and if everyone 219redials on the same schedule the line stays jammed, so you add a random 220pause (**jitter**). 221 222**Worked example.** A service times out twice and then answers. With a base 223delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt 224succeeds. With a cap of 10 s, the waits never exceed 10 s however many 225attempts there are. A request that's *invalid* (a `ValueError` here) fails 226the same way every time, so it's raised immediately: retrying only adds 227load. 228 229$$ 230d_k = \min\bigl(D_{\max},\; b \cdot 2^{k}\bigr), \qquad k = 0, 1, 2, \ldots 231$$ 232 233**Symbols** 234 235| Symbol | Meaning here | Worked example | 236|---|---|---| 237| $d_k$ | wait before retry number $k+1$ | $d_0$ = 0.1 s, $d_1$ = 0.2 s | 238| $k$ | how many retries have already happened | 0, 1 | 239| $b$ | base delay | 0.1 s | 240| $2^{k}$ | doubles with every retry | 1, 2, 4, … | 241| $D_{\max}$ | the cap on any single wait | 10 s | 242 243**In words:** each wait doubles the previous one, starting from the base 244delay, but never exceeds the cap. 245 246**On the worked example:** $d_0$ = min(10, 0.1 × 1) = 0.1 s and 247$d_1$ = min(10, 0.1 × 2) = 0.2 s. 248 249**In Python:** 250 251```python 252b, D_max = 0.1, 10 253# d_0 and d_1 254[min(D_max, b * 2 ** k) for k in range(2)] # → [0.1, 0.2] 255``` 256 257```mermaid 258sequenceDiagram 259 participant A as Agent tool 260 participant S as Upstream API 261 A->>S: request 262 S--xA: timeout 263 Note over A: wait 0.1 s 264 A->>S: retry 1 265 S--xA: timeout 266 Note over A: wait 0.2 s 267 A->>S: retry 2 268 S-->>A: 200 OK 269``` 270 271**Reading it:** time runs downward. Each failure is followed by a longer 272pause, which gives an overloaded service room to recover instead of being 273hammered. Writes must be **idempotent** (safe to repeat: sending the same 274request twice has the same effect as once, usually via an idempotency key) 275before you retry them, or a retried "create order" makes two orders. 276 277 278 279**Reading it:** the x-axis is the attempt number; the black line is the 280capped exponential delay. The dots are five clients using *full jitter* 281(each waits a random time between zero and the capped delay): instead of 282all retrying at the same moment, their retries spread out, which is what 283lets a recovering service actually recover. 284 285**In code:** `backoff_delays` lists the capped waits $d_k$; 286`retry_with_backoff` calls a function, retries only the exceptions in 287`RETRYABLE` with those waits (optionally with full jitter), and raises 288anything else at once. 289 290## Brittle integrations: circuit breakers 291 292**Everyday picture.** The breaker in your home's fuse box. When a circuit 293keeps shorting, it trips and cuts the power, instead of letting the wire 294overheat. After a while you flip it back on to test; if it trips again, it 295stays off. 296 297A **circuit breaker** wraps calls to a dependency. After several 298consecutive failures it **opens**: calls fail immediately without touching 299the struggling service (fail fast), so the agent can fall back (use a cache, 300tell the user, hand off to a person) instead of hanging on timeouts. After a 301cool-down it goes **half-open** and lets one trial call through: success 302closes it, failure opens it again. 303 304**Worked example.** Threshold 3, cool-down 30 s. Three timeouts in a row: 305open. A fourth call at 5 s fails instantly with `CircuitOpen` and the 306service isn't called at all. At 31 s, one trial call goes through; it 307succeeds, so the breaker closes. 308 309```mermaid 310stateDiagram-v2 311 [*] --> Closed 312 Closed --> Closed: success (reset the failure count) 313 Closed --> Open: 3 failures in a row 314 Open --> Open: calls rejected instantly 315 Open --> HalfOpen: cool-down elapsed 316 HalfOpen --> Closed: trial call succeeds 317 HalfOpen --> Open: trial call fails 318``` 319 320**Reading it:** in the normal state (Closed) calls flow through and failures 321are counted. Open is the protective state: nothing reaches the dependency. 322Half-open is a single, cautious test. The breaker turns a slow, cascading 323failure (every request waiting 30 s on a dead service) into a fast, 324contained one. 325 326 327 328**Reading it:** the shaded band is when the upstream service is down; 329each marker is one request. Before the outage, calls succeed (green). The 330first three failures (red) open the breaker; after that, requests fail fast 331(grey) without reaching the service, apart from one trial call per 332cool-down. Once the service is back, the next trial succeeds and traffic 333resumes. The service received a handful of calls during the outage instead 334of all of them. 335 336**In code:** `CircuitBreaker.call` is the state diagram: it raises 337`CircuitOpen` while open, lets one trial call through after the cool-down, 338and closes on success. `simulate_outage` replays the outage in the figure. 339 340**Contract tests** (automated checks that an external API still accepts 341and returns what your tool expects) catch the third kind of brittleness, 342upstream API changes, before users do. 343 344## Loops and runaway 345 346**Everyday picture.** A satnav that keeps rerouting you around the same 347block. The fix isn't a better map; it's noticing that you've passed the same 348corner three times. 349 350**Worked example.** `search("vpn")`, `search("vpn")`, `search("vpn")`: the 351same tool with the same arguments three times in a row. Nothing new can 352come back, so the agent is looping. `search("vpn")`, `search("vpn error")`, 353`search("ERR-4012")` is the same tool refining its query: progress. 354Combine loop detection with hard step and token budgets 355(`primer.agents.cost.TaskBudget`) so a confused agent stops and hands off 356instead of burning money. 357 358**In code:** `is_looping` reports whether the last few calls were the same 359tool with identical arguments. 360 361## In 20 seconds 362- Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60% 363 of the time. Shorten chains and verify at key steps. 364- Retry transient failures with capped exponential backoff and jitter; never 365 retry invalid requests; make writes idempotent first. 366- Put circuit breakers around dependencies so outages fail fast instead of 367 cascading. 368- Detect loops (same call, same arguments) and enforce step and token 369 budgets. 370- Most failures are system failures (retrieval, tools, data, evals, 371 adoption), not model failures, and each has a known fix. 372 373## Self-test questions 374 375**Your agent works in demos and fails half the time in production. Where do 376you look first?** 377At the length of real tasks versus demo tasks (compounding error), and at 378traces for the first failing step: retrieval misses, wrong tools or 379arguments, tool errors from integrations, context growth. Then fix 380per-step reliability where it's lowest, add verification after key steps, 381and put the failing cases into the eval set. 382 383**Why add jitter to retries?** 384Without it, every client that failed at the same moment retries at the same 385moment, producing synchronized waves of load that keep an overloaded 386service down. Random delays spread the retries out. 387 388**When should you not retry?** 389When the error can't go away on its own: invalid input, permission denied, 390not found. And never retry a non-idempotent write without an idempotency 391key, or you'll duplicate its effect. 392 393**What's the difference between a retry and a circuit breaker?** 394A retry handles a blip for one request. A breaker handles an outage across 395many requests: it stops sending traffic to a failing dependency, fails fast 396so callers can fall back, and probes for recovery. 397 398## The papers behind this lesson 399 400No single founding paper; these are engineering patterns. The canonical 401write-ups are Michael Nygard's *Release It!* (which named the circuit 402breaker) and the AWS Builders' Library article on backoff and jitter, both 403linked below. 404 405## Further reading 406- Martin Fowler, *CircuitBreaker*: https://martinfowler.com/bliki/CircuitBreaker.html 407- AWS Builders' Library, *Timeouts, retries, and backoff with jitter*: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/ 408- Anthropic, *Building effective agents*: https://www.anthropic.com/engineering/building-effective-agents 409""" 410 411from __future__ import annotations 412 413import json 414import math 415import time 416from dataclasses import dataclass 417from typing import Any, Callable 418 419import numpy as np 420 421from primer._show import banner, say, table, takeaway 422 423# --------------------------------------------------------------------------- 424# 1. The catalogue 425# --------------------------------------------------------------------------- 426 427 428@dataclass(frozen=True) 429class Failure: 430 name: str 431 symptom: str 432 fix: str 433 module: str 434 435 436FAILURES: list[Failure] = [ 437 Failure("Compounding error", "Long tasks fail far more than short ones", "Fewer steps, verify key steps, checkpoints", "primer.agents.planning"), 438 Failure("Bad retrieval", "Confident wrong answers", "Hybrid search, reranking, retrieval evals", "primer.agents.rag"), 439 Failure("Ambiguous tools", "Wrong tool or bad arguments", "Precise descriptions, fewer tools, validation", "primer.agents.tools"), 440 Failure("Context rot", "Quality drops as sessions grow", "Summarize, trim, restart with a state handoff", "primer.agents.context"), 441 Failure("Loops and runaway", "Repeated calls, cost spikes", "Step budgets, loop detection, stop conditions", "primer.agents.agent_loop"), 442 Failure("No evals", "Regressions ship silently", "Golden sets in CI, online monitoring", "primer.agents.evals"), 443 Failure("Prompt injection", "Agent follows instructions found in data", "Untrusted-content boundaries, privilege separation", "primer.agents.guardrails"), 444 Failure("Messy enterprise data", "Garbled parsing, permission gaps", "Invest in ingestion, permission-aware retrieval", "primer.agents.rag"), 445 Failure("Brittle integrations", "Timeouts, auth expiry, API changes", "Retries with backoff, circuit breakers, contract tests", "primer.agents.failures"), 446 Failure("No adoption", "Works, but nobody uses it", "Design with users, show sources, easy human handoff", "primer.agents.deployment"), 447] 448 449# --------------------------------------------------------------------------- 450# 2. Compounding error 451# --------------------------------------------------------------------------- 452 453 454def chain_success(p: float, n: int) -> float: 455 """Chance that n independent steps, each succeeding with probability p, all succeed.""" 456 return p**n 457 458 459def max_steps_for(p: float, target: float) -> int: 460 """The most steps you can chain at per-step reliability p and still reach `target`.""" 461 # p**n >= target <=> n <= log(target) / log(p) (both logs are negative) 462 return math.floor(math.log(target) / math.log(p) + 1e-12) 463 464 465# --------------------------------------------------------------------------- 466# 3. Retries with exponential backoff 467# --------------------------------------------------------------------------- 468 469RETRYABLE: tuple[type[Exception], ...] = (TimeoutError, ConnectionError) 470 471 472def backoff_delays(attempts: int, base_delay: float, max_delay: float) -> list[float]: 473 """The capped exponential wait before each retry (no jitter).""" 474 return [min(max_delay, base_delay * 2**k) for k in range(attempts - 1)] 475 476 477def retry_with_backoff(fn: Callable[[], Any], max_attempts: int = 5, base_delay: float = 0.1, max_delay: float = 10.0, 478 jitter: np.random.Generator | None = None, sleep: Callable[[float], None] = time.sleep, 479 retryable: tuple[type[Exception], ...] = RETRYABLE) -> Any: 480 """Call fn, retrying transient failures with capped exponential backoff. 481 482 Pass a seeded `np.random.default_rng()` as `jitter` for "full jitter": 483 wait a uniform random time between 0 and the capped delay. 484 """ 485 for attempt in range(max_attempts): 486 try: 487 return fn() 488 except retryable: 489 if attempt == max_attempts - 1: 490 raise 491 delay = min(max_delay, base_delay * 2**attempt) 492 sleep(float(jitter.uniform(0, delay)) if jitter is not None else delay) 493 raise AssertionError("unreachable") 494 495 496# --------------------------------------------------------------------------- 497# 4. Circuit breaker 498# --------------------------------------------------------------------------- 499 500 501class CircuitOpen(RuntimeError): 502 """Raised instead of calling a dependency whose breaker is open.""" 503 504 505class CircuitBreaker: 506 def __init__(self, failure_threshold: int = 3, reset_timeout: float = 30.0, clock: Callable[[], float] = time.monotonic): 507 self.threshold, self.reset_timeout, self.clock = failure_threshold, reset_timeout, clock 508 self.state = "closed" 509 self.failures = 0 510 self.opened_at = 0.0 511 512 def call(self, fn: Callable[[], Any]) -> Any: 513 if self.state == "open": 514 if self.clock() - self.opened_at < self.reset_timeout: 515 raise CircuitOpen("dependency unavailable; failing fast") 516 self.state = "half_open" # cool-down over: allow one trial call 517 try: 518 result = fn() 519 except Exception: 520 self.failures += 1 521 if self.state == "half_open" or self.failures >= self.threshold: 522 self.state, self.opened_at = "open", self.clock() 523 raise 524 self.state, self.failures = "closed", 0 525 return result 526 527 528def simulate_outage(outage: tuple[float, float] = (10.0, 60.0), every: float = 2.0, until: float = 100.0) -> list[dict[str, Any]]: 529 """Requests every `every` seconds through a breaker while the service is down during `outage`.""" 530 now = [0.0] 531 breaker = CircuitBreaker(failure_threshold=3, reset_timeout=15.0, clock=lambda: now[0]) 532 events = [] 533 534 def service(): 535 if outage[0] <= now[0] < outage[1]: 536 raise TimeoutError("down") 537 return "ok" 538 539 t = 0.0 540 while t < until: 541 now[0] = t 542 try: 543 breaker.call(service) 544 outcome = "ok" 545 except CircuitOpen: 546 outcome = "fast-fail" 547 except TimeoutError: 548 outcome = "failed" 549 events.append({"t": t, "outcome": outcome, "state": breaker.state}) 550 t += every 551 return events 552 553 554# --------------------------------------------------------------------------- 555# 5. Loop detection 556# --------------------------------------------------------------------------- 557 558 559def is_looping(calls: list[tuple[str, dict[str, Any]]], repeats: int = 3) -> bool: 560 """True if the last `repeats` calls are the same tool with identical arguments.""" 561 if len(calls) < repeats: 562 return False 563 keys = [(name, json.dumps(args, sort_keys=True)) for name, args in calls[-repeats:]] 564 return len(set(keys)) == 1 565 566 567# --------------------------------------------------------------------------- 568# Figures and demo 569# --------------------------------------------------------------------------- 570 571 572def figures() -> dict[str, Any]: 573 import matplotlib 574 575 matplotlib.use("Agg") 576 import matplotlib.pyplot as plt 577 578 figs: dict[str, Any] = {} 579 580 n = np.arange(1, 31) 581 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 582 for p, color in [(0.99, "#55a868"), (0.95, "#4c72b0"), (0.90, "#c44e52")]: 583 ax.plot(n, [chain_success(p, k) for k in n], color=color, label=f"{p:.0%} per step") 584 for k in (10, 20): 585 ax.plot(k, chain_success(0.95, k), "o", color="#4c72b0") 586 ax.annotate(f"{chain_success(0.95, k):.0%}", (k, chain_success(0.95, k)), textcoords="offset points", xytext=(5, 5), fontsize=8) 587 ax.set_xlabel("steps that must all succeed") 588 ax.set_ylabel("chance the whole task succeeds") 589 ax.set_ylim(0, 1.02) 590 ax.set_title("Compounding error") 591 ax.legend() 592 fig.tight_layout() 593 figs["compounding"] = fig 594 595 attempts = 8 596 base = backoff_delays(attempts, 0.5, 10.0) 597 rng = np.random.default_rng(0) 598 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 599 ax.plot(range(1, attempts), base, "k-o", label="capped exponential (no jitter)") 600 for c in range(5): 601 ax.scatter(np.arange(1, attempts) + (c - 2) * 0.06, [rng.uniform(0, d) for d in base], s=18, alpha=0.8, 602 label="five clients with full jitter" if c == 0 else None, color="#4c72b0") 603 ax.set_xlabel("retry number") 604 ax.set_ylabel("seconds to wait") 605 ax.set_title("Backoff: base 0.5 s, doubling, capped at 10 s") 606 ax.legend(fontsize=8) 607 fig.tight_layout() 608 figs["backoff"] = fig 609 610 ev = simulate_outage() 611 fig, ax = plt.subplots(figsize=(7, 2.8)) 612 ax.axvspan(10, 60, color="#f2d0d0", label="service down") 613 colors = {"ok": "#55a868", "failed": "#c44e52", "fast-fail": "#8c8c8c"} 614 for kind in colors: 615 pts = [e["t"] for e in ev if e["outcome"] == kind] 616 ax.scatter(pts, [1] * len(pts), color=colors[kind], s=40, label=kind.replace("fast-fail", "failed fast (breaker open)")) 617 ax.set_yticks([]) 618 ax.set_xlabel("seconds") 619 ax.set_title("Circuit breaker (threshold 3, cool-down 15 s) during an outage") 620 ax.legend(fontsize=7, loc="upper center", bbox_to_anchor=(0.5, -0.35), ncol=4, frameon=False) 621 fig.tight_layout() 622 figs["breaker"] = fig 623 return figs 624 625 626def demo() -> None: 627 banner("1. The catalogue") 628 table(["failure", "symptom", "fix", "lesson"], [(f.name, f.symptom, f.fix, f.module) for f in FAILURES]) 629 630 banner("2. Compounding error") 631 table(["steps", "at 99%", "at 95%", "at 90%"], [(k, chain_success(0.99, k), chain_success(0.95, k), chain_success(0.90, k)) for k in (1, 5, 10, 20)], floatfmt=".3f") 632 print(f"at 95% per step, the most steps that keep success at 90% or better: {max_steps_for(0.95, 0.90)}") 633 print() 634 635 banner("3. Retries with backoff") 636 waits: list[float] = [] 637 calls = {"n": 0} 638 639 def flaky(): 640 calls["n"] += 1 641 # Two timeouts, then an answer: the worked example in the lesson. 642 if calls["n"] <= 2: 643 raise TimeoutError("upstream timed out") 644 return "ok" 645 646 print("result:", retry_with_backoff(flaky, base_delay=0.1, sleep=waits.append), "after waits", waits) 647 print() 648 649 banner("4. Circuit breaker during an outage") 650 ev = simulate_outage() 651 reached = sum(e["outcome"] == "failed" for e in ev) 652 fast = sum(e["outcome"] == "fast-fail" for e in ev) 653 print(f"during the outage the service was hit {reached} times; {fast} calls failed fast instead") 654 print() 655 656 banner("5. Loop detection") 657 print(is_looping([("search", {"q": "vpn"})] * 3), is_looping([("search", {"q": "vpn"}), ("search", {"q": "vpn error"}), ("search", {"q": "ERR-4012"})])) 658 print() 659 say("Same tool, same arguments, three times: nothing new can come back. Stop, and hand off.") 660 takeaway("Most agent failures are system failures, and each one has a known fix. Build the fix, then test for it.") 661 662 663if __name__ == "__main__": 664 demo()
455def chain_success(p: float, n: int) -> float: 456 """Chance that n independent steps, each succeeding with probability p, all succeed.""" 457 return p**n
Chance that n independent steps, each succeeding with probability p, all succeed.
460def max_steps_for(p: float, target: float) -> int: 461 """The most steps you can chain at per-step reliability p and still reach `target`.""" 462 # p**n >= target <=> n <= log(target) / log(p) (both logs are negative) 463 return math.floor(math.log(target) / math.log(p) + 1e-12)
The most steps you can chain at per-step reliability p and still reach target.
473def backoff_delays(attempts: int, base_delay: float, max_delay: float) -> list[float]: 474 """The capped exponential wait before each retry (no jitter).""" 475 return [min(max_delay, base_delay * 2**k) for k in range(attempts - 1)]
The capped exponential wait before each retry (no jitter).
478def retry_with_backoff(fn: Callable[[], Any], max_attempts: int = 5, base_delay: float = 0.1, max_delay: float = 10.0, 479 jitter: np.random.Generator | None = None, sleep: Callable[[float], None] = time.sleep, 480 retryable: tuple[type[Exception], ...] = RETRYABLE) -> Any: 481 """Call fn, retrying transient failures with capped exponential backoff. 482 483 Pass a seeded `np.random.default_rng()` as `jitter` for "full jitter": 484 wait a uniform random time between 0 and the capped delay. 485 """ 486 for attempt in range(max_attempts): 487 try: 488 return fn() 489 except retryable: 490 if attempt == max_attempts - 1: 491 raise 492 delay = min(max_delay, base_delay * 2**attempt) 493 sleep(float(jitter.uniform(0, delay)) if jitter is not None else delay) 494 raise AssertionError("unreachable")
Call fn, retrying transient failures with capped exponential backoff.
Pass a seeded np.random.default_rng() as jitter for "full jitter":
wait a uniform random time between 0 and the capped delay.
502class CircuitOpen(RuntimeError): 503 """Raised instead of calling a dependency whose breaker is open."""
Raised instead of calling a dependency whose breaker is open.
506class CircuitBreaker: 507 def __init__(self, failure_threshold: int = 3, reset_timeout: float = 30.0, clock: Callable[[], float] = time.monotonic): 508 self.threshold, self.reset_timeout, self.clock = failure_threshold, reset_timeout, clock 509 self.state = "closed" 510 self.failures = 0 511 self.opened_at = 0.0 512 513 def call(self, fn: Callable[[], Any]) -> Any: 514 if self.state == "open": 515 if self.clock() - self.opened_at < self.reset_timeout: 516 raise CircuitOpen("dependency unavailable; failing fast") 517 self.state = "half_open" # cool-down over: allow one trial call 518 try: 519 result = fn() 520 except Exception: 521 self.failures += 1 522 if self.state == "half_open" or self.failures >= self.threshold: 523 self.state, self.opened_at = "open", self.clock() 524 raise 525 self.state, self.failures = "closed", 0 526 return result
513 def call(self, fn: Callable[[], Any]) -> Any: 514 if self.state == "open": 515 if self.clock() - self.opened_at < self.reset_timeout: 516 raise CircuitOpen("dependency unavailable; failing fast") 517 self.state = "half_open" # cool-down over: allow one trial call 518 try: 519 result = fn() 520 except Exception: 521 self.failures += 1 522 if self.state == "half_open" or self.failures >= self.threshold: 523 self.state, self.opened_at = "open", self.clock() 524 raise 525 self.state, self.failures = "closed", 0 526 return result
529def simulate_outage(outage: tuple[float, float] = (10.0, 60.0), every: float = 2.0, until: float = 100.0) -> list[dict[str, Any]]: 530 """Requests every `every` seconds through a breaker while the service is down during `outage`.""" 531 now = [0.0] 532 breaker = CircuitBreaker(failure_threshold=3, reset_timeout=15.0, clock=lambda: now[0]) 533 events = [] 534 535 def service(): 536 if outage[0] <= now[0] < outage[1]: 537 raise TimeoutError("down") 538 return "ok" 539 540 t = 0.0 541 while t < until: 542 now[0] = t 543 try: 544 breaker.call(service) 545 outcome = "ok" 546 except CircuitOpen: 547 outcome = "fast-fail" 548 except TimeoutError: 549 outcome = "failed" 550 events.append({"t": t, "outcome": outcome, "state": breaker.state}) 551 t += every 552 return events
Requests every every seconds through a breaker while the service is down during outage.
560def is_looping(calls: list[tuple[str, dict[str, Any]]], repeats: int = 3) -> bool: 561 """True if the last `repeats` calls are the same tool with identical arguments.""" 562 if len(calls) < repeats: 563 return False 564 keys = [(name, json.dumps(args, sort_keys=True)) for name, args in calls[-repeats:]] 565 return len(set(keys)) == 1
True if the last repeats calls are the same tool with identical arguments.
573def figures() -> dict[str, Any]: 574 import matplotlib 575 576 matplotlib.use("Agg") 577 import matplotlib.pyplot as plt 578 579 figs: dict[str, Any] = {} 580 581 n = np.arange(1, 31) 582 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 583 for p, color in [(0.99, "#55a868"), (0.95, "#4c72b0"), (0.90, "#c44e52")]: 584 ax.plot(n, [chain_success(p, k) for k in n], color=color, label=f"{p:.0%} per step") 585 for k in (10, 20): 586 ax.plot(k, chain_success(0.95, k), "o", color="#4c72b0") 587 ax.annotate(f"{chain_success(0.95, k):.0%}", (k, chain_success(0.95, k)), textcoords="offset points", xytext=(5, 5), fontsize=8) 588 ax.set_xlabel("steps that must all succeed") 589 ax.set_ylabel("chance the whole task succeeds") 590 ax.set_ylim(0, 1.02) 591 ax.set_title("Compounding error") 592 ax.legend() 593 fig.tight_layout() 594 figs["compounding"] = fig 595 596 attempts = 8 597 base = backoff_delays(attempts, 0.5, 10.0) 598 rng = np.random.default_rng(0) 599 fig, ax = plt.subplots(figsize=(6.5, 3.5)) 600 ax.plot(range(1, attempts), base, "k-o", label="capped exponential (no jitter)") 601 for c in range(5): 602 ax.scatter(np.arange(1, attempts) + (c - 2) * 0.06, [rng.uniform(0, d) for d in base], s=18, alpha=0.8, 603 label="five clients with full jitter" if c == 0 else None, color="#4c72b0") 604 ax.set_xlabel("retry number") 605 ax.set_ylabel("seconds to wait") 606 ax.set_title("Backoff: base 0.5 s, doubling, capped at 10 s") 607 ax.legend(fontsize=8) 608 fig.tight_layout() 609 figs["backoff"] = fig 610 611 ev = simulate_outage() 612 fig, ax = plt.subplots(figsize=(7, 2.8)) 613 ax.axvspan(10, 60, color="#f2d0d0", label="service down") 614 colors = {"ok": "#55a868", "failed": "#c44e52", "fast-fail": "#8c8c8c"} 615 for kind in colors: 616 pts = [e["t"] for e in ev if e["outcome"] == kind] 617 ax.scatter(pts, [1] * len(pts), color=colors[kind], s=40, label=kind.replace("fast-fail", "failed fast (breaker open)")) 618 ax.set_yticks([]) 619 ax.set_xlabel("seconds") 620 ax.set_title("Circuit breaker (threshold 3, cool-down 15 s) during an outage") 621 ax.legend(fontsize=7, loc="upper center", bbox_to_anchor=(0.5, -0.35), ncol=4, frameon=False) 622 fig.tight_layout() 623 figs["breaker"] = fig 624 return figs
627def demo() -> None: 628 banner("1. The catalogue") 629 table(["failure", "symptom", "fix", "lesson"], [(f.name, f.symptom, f.fix, f.module) for f in FAILURES]) 630 631 banner("2. Compounding error") 632 table(["steps", "at 99%", "at 95%", "at 90%"], [(k, chain_success(0.99, k), chain_success(0.95, k), chain_success(0.90, k)) for k in (1, 5, 10, 20)], floatfmt=".3f") 633 print(f"at 95% per step, the most steps that keep success at 90% or better: {max_steps_for(0.95, 0.90)}") 634 print() 635 636 banner("3. Retries with backoff") 637 waits: list[float] = [] 638 calls = {"n": 0} 639 640 def flaky(): 641 calls["n"] += 1 642 # Two timeouts, then an answer: the worked example in the lesson. 643 if calls["n"] <= 2: 644 raise TimeoutError("upstream timed out") 645 return "ok" 646 647 print("result:", retry_with_backoff(flaky, base_delay=0.1, sleep=waits.append), "after waits", waits) 648 print() 649 650 banner("4. Circuit breaker during an outage") 651 ev = simulate_outage() 652 reached = sum(e["outcome"] == "failed" for e in ev) 653 fast = sum(e["outcome"] == "fast-fail" for e in ev) 654 print(f"during the outage the service was hit {reached} times; {fast} calls failed fast instead") 655 print() 656 657 banner("5. Loop detection") 658 print(is_looping([("search", {"q": "vpn"})] * 3), is_looping([("search", {"q": "vpn"}), ("search", {"q": "vpn error"}), ("search", {"q": "ERR-4012"})])) 659 print() 660 say("Same tool, same arguments, three times: nothing new can come back. Stop, and hand off.") 661 takeaway("Most agent failures are system failures, and each one has a known fix. Build the fix, then test for it.")