primer.agents.failures

Why the hard ones fail: a catalogue of agent failures, and the fix for each

Run: python -m primer.agents.failures

This lesson builds on the agent loop from primer.agents.agent_loop, on traces from primer.agents.observability and on the golden set from primer.agents.evals.

Level 1: The practitioner's guide

In one sentence. Most agent failures in production are system failures with known fixes (too many steps, bad retrieval, ambiguous tools, a context that rots, loops, no evals, injected instructions, messy data, brittle integrations, no adoption), and this lesson is the catalogue: symptom, fix, and where the fix is built.

When you need it. When the agent works in demos and fails in production, and nobody can say why. The tell is the gap between the two: demos are three steps long and real tasks are twenty. This lesson's arithmetic explains the gap on its own. If each step succeeds 95% of the time, a ten-step task succeeds 60% of the time and a twenty-step task 36%; at 90% per step, twenty steps succeed 12% of the time; at 99%, 82%. At 95% per step, the most steps you can chain and still succeed nine times in ten is two. Nothing about the model changed between the demo and production; the exponent did. You need this catalogue before the first incident, as a checklist, and after every incident, to name what happened. You don't need the integration patterns (retries, breakers) for a tool that calls nothing outside your process, and you don't need loop detection for a workflow with a fixed number of steps.

Your options. Six defences, from the ones that protect one request to the ones that protect the whole system:

Option What it does What it guarantees What it costs Where it lives
Retries with capped exponential backoff and jitter Wait 0.1 s, then 0.2 s, doubling to a cap, with a random spread so clients don't retry in waves A transient blip (a timeout, a rate limit) doesn't fail the task Latency on the retried request; a duplicate write unless the write is idempotent Around each outside call
Circuit breaker After a run of failures, calls fail instantly for a cool-down, then one trial call probes for recovery An outage fails fast and stays contained instead of tying every request up in timeouts (in this lesson's 50-second outage the dead service is called 5 times and 21 requests fail fast) A fallback path for when the breaker is open Around each dependency
Loop detection and budgets Stop when the same tool is called with the same arguments three times, or when a step or token budget runs out A confused agent stops and hands off instead of burning money all night A false stop now and then; a handoff path The agent loop
Fewer steps, verification, checkpoints Shorten the chain, check a step's result where it happens and retry only that step, save state to resume from the middle Attacks the exponent directly: errors stop travelling down the chain Design effort; a check after each key step costs a call The task design
Contract tests Automated checks that an outside API still accepts and returns what your tool expects An upstream change is caught by a test, not a user Tests to write and keep current Your test suite
Evals and traces A golden set in CI and a trace of every run Regressions don't ship silently, and every failure can be replayed to its first bad step The two lessons before this one Your pipeline and your monitoring

How to choose. Read the symptom first; the catalogue maps each one to a fix.

  • Long tasks fail and short ones don't: compounding error. Cut steps, verify after the key ones, checkpoint, and put a person at the critical points.
  • Timeouts and expired credentials: brittle integrations. Retries with backoff for blips, a breaker for outages, contract tests for changes. Never retry an invalid request (it fails the same way every time), and never retry a write that isn't idempotent.
  • The same call repeated, cost spiking: loops. Detect the repeat and set hard budgets.
  • Confident wrong answers: bad retrieval, or messy source data. Fix search and ingestion before touching the prompt.
  • Wrong tool or bad arguments: ambiguous tools. Fewer tools, precise descriptions, validated inputs.
  • The agent obeys text it found in a document: prompt injection. Boundaries around untrusted content and privilege separation.
  • It works and nobody uses it: adoption. Build with the people it serves.
  • Whatever the symptom, look at the trace for the first failing step before changing anything, and add the case to the golden set.

What it costs. Retries cost time on the slow path: the lesson's waits are 0.1 s and 0.2 s before the third attempt succeeds, capped at 10 s however many attempts there are, and with full jitter each client waits a random time up to that cap. A breaker costs almost nothing while closed and saves a great deal while open: every request that would have waited 30 s on a dead service fails at once instead. Loop detection is a comparison of the last few calls. Verification after a step costs a model call per check, which is why it goes after the key steps, not all of them. The expensive fix is the design one, shortening chains, and it is also the one with the largest payoff: raising per-step reliability from 95% to 99% takes a twenty-step task from 36% to 82%. The lesson's figures draw each of these curves so the trade can be read off, not guessed.

What breaks.

  • Retrying the invalid. A bad request retried five times is five times the load and the same error. Retry only the exceptions that can clear on their own.
  • Retrying a write without an idempotency key. "Create order" twice is two orders.
  • Retries in waves. Every client that failed at the same moment retries at the same moment, and the recovering service goes down again. Jitter.
  • No breaker. One dead dependency, and every request waits out a full timeout; the whole agent looks down.
  • A breaker with no fallback. Failing fast is only useful if the agent can do something else: use a cache, tell the user, hand off.
  • Loop detection that only counts. Three calls of the same tool with refining queries is progress, not a loop. Compare arguments, not names.
  • Fixing the prompt for a system fault. A retrieval miss, a tool error or a parsing bug looks like a model mistake from the answer alone. The trace shows which it was.

In the wild. The circuit breaker was named by Michael Nygard in Release It! and written up by Martin Fowler with the three states this lesson builds (closed, open, half-open); the retry guidance, capped exponential backoff with jitter, is the AWS Builders' Library article the lesson links. Libraries package both: in Python, tenacity gives a retry decorator with exponential and random exponential waits, stop conditions by attempts or elapsed time, and retry only on chosen exception types, the same policy as this lesson's retry_with_backoff. Anthropic's Building effective agents draws the same conclusions from the other direction: start with simple prompts and evaluation, add multi-step agents only where simpler designs fall short, add complexity only when it demonstrably improves outcomes, and have the agent pause for human feedback at checkpoints; its emphasis on tool design as an interface problem is the fix for the ambiguous-tools row. The seven rows this lesson doesn't build point to the lessons that do.

Go deeper. Level 2 builds three of the fixes in plain code: the compounding-error formula and its inverse (how many steps a target allows), retries with capped backoff and full jitter that raise invalid requests at once, a circuit breaker as a three-state machine replayed through an outage, and a loop detector that compares arguments. If you only needed to name the failure and find its fix, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

Pilots learn from a catalogue of accidents: each entry names what went wrong, how it showed up in the cockpit, and the checklist item that now prevents it. This lesson is that catalogue for agents: ten failures that keep recurring in production systems, each with its symptom, its fix, and a pointer to the lesson that builds the fix. Three of them get their fix built right here: compounding error (arithmetic), brittle integrations (retries with backoff, and circuit breakers) and runaway loops (loop detection).

Failure Symptom Fix Lesson
Compounding error Long tasks fail far more often than short ones Fewer steps, checks after key steps, checkpoints primer.agents.planning
Bad retrieval Confident, wrong answers Hybrid search, reranking, retrieval evals primer.agents.rag
Ambiguous tools Wrong tool, or bad arguments Precise descriptions, fewer tools, validation primer.agents.tools
Context rot Quality drops as a session grows Summarize, trim, restart with a state handoff primer.agents.context
Loops and runaway The same call repeated; cost spikes Step budgets, loop detection, stop conditions primer.agents.agent_loop
No evals Regressions ship silently Golden sets in CI, online monitoring primer.agents.evals
Prompt injection The agent obeys instructions found in data Untrusted-content boundaries, privilege separation primer.agents.guardrails
Messy enterprise data Garbled tables, missing permissions Invest in parsing; permission-aware retrieval primer.agents.rag
Brittle integrations Timeouts, expired credentials, API changes Retries with backoff, circuit breakers, contract tests primer.agents.failures
No adoption It works, and nobody uses it Build with users, show sources, easy human handoff primer.agents.deployment

In code: FAILURES is this table as data, one Failure record per row.

Compounding error

Everyday picture. A relay race where each baton pass succeeds 95% of the time. One pass is nearly safe; ten passes in a row drop the baton more often than you'd think.

Worked example. If each step of an agent's task succeeds 95% of the time, independently:

Steps Chance every step succeeds
1 0.95
2 0.95 × 0.95 = 0.9025
10 0.95¹⁰ ≈ 0.599
20 0.95²⁰ ≈ 0.358
Level 3: the formula and its symbols

$$ P(\text{task succeeds}) = p^{\,n} $$

Symbols

Symbol Meaning here Worked example
$p$ chance a single step succeeds 0.95
$n$ number of steps that must all succeed 10
$p^{\,n}$ $p$ multiplied by itself $n$ times 0.599

In words: the chance that every step succeeds is the per-step chance multiplied together once per step.

On the worked example: 0.95 multiplied by itself 10 times is 0.599, so a ten-step task at 95% per step fails four times in ten.

Level 3: in Python

In Python:

p, n = 0.95, 10
# p multiplied by itself n times
round(p ** n, 3)  # → 0.599
# fails about four times in ten
round(1 - p ** n, 1)  # → 0.4
flowchart LR S1[Step 1<br/>95%] --> S2[Step 2<br/>95%] --> S3[...] --> S10[Step 10<br/>95%] --> D[Done:<br/>60% of the time] S2 -.->|verify, retry<br/>just this step| S2

Reading it: every arrow is another chance to drop the baton, so success shrinks with each step. The dotted loop is the remedy: check a step's result where it happens and retry only that step, so an error doesn't travel down the chain. Fewer steps, verification after key steps, checkpoints to resume from the middle, and human review at critical points all attack the same exponent.

Twenty steps succeed 82% of the time at 99% per step but only 36% at 95% and 12% at 90%

Reading it: the x-axis is the number of steps; each curve is a per-step reliability. At 99% per step a 20-step task still succeeds 82% of the time; at 95% it's 36%; at 90%, 12%. Small gains in per-step reliability are worth far more than they look, and demos with three steps say little about tasks with twenty.

In code: chain_success is $p^{\,n}$; max_steps_for runs it backwards, returning the most steps you can chain and still reach a target success rate.

Brittle integrations: retries with backoff

Everyday picture. Calling a busy phone line: you don't redial every second. You wait a little, then longer, then longer still, and if everyone redials on the same schedule the line stays jammed, so you add a random pause (jitter).

Worked example. A service times out twice and then answers. With a base delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt succeeds. With a cap of 10 s, the waits never exceed 10 s however many attempts there are. A request that's invalid (a ValueError here) fails the same way every time, so it's raised immediately: retrying only adds load.

Level 3: the formula and its symbols

$$ d_k = \min\bigl(D_{\max},\; b \cdot 2^{k}\bigr), \qquad k = 0, 1, 2, \ldots $$

Symbols

Symbol Meaning here Worked example
$d_k$ wait before retry number $k+1$ $d_0$ = 0.1 s, $d_1$ = 0.2 s
$k$ how many retries have already happened 0, 1
$b$ base delay 0.1 s
$2^{k}$ doubles with every retry 1, 2, 4, …
$D_{\max}$ the cap on any single wait 10 s

In words: each wait doubles the previous one, starting from the base delay, but never exceeds the cap.

On the worked example: $d_0$ = min(10, 0.1 × 1) = 0.1 s and $d_1$ = min(10, 0.1 × 2) = 0.2 s.

Level 3: in Python

In Python:

b, D_max = 0.1, 10
# d_0 and d_1
[min(D_max, b * 2 ** k) for k in range(2)]  # → [0.1, 0.2]
sequenceDiagram participant A as Agent tool participant S as Upstream API A->>S: request S--xA: timeout Note over A: wait 0.1 s A->>S: retry 1 S--xA: timeout Note over A: wait 0.2 s A->>S: retry 2 S-->>A: 200 OK

Reading it: time runs downward. Each failure is followed by a longer pause, which gives an overloaded service room to recover instead of being hammered. Writes must be idempotent (safe to repeat: sending the same request twice has the same effect as once, usually via an idempotency key) before you retry them, or a retried "create order" makes two orders.

The delay doubles each attempt until it hits the 10-second cap, and jitter scatters five clients' retries across the whole range below it

Reading it: the x-axis is the attempt number; the black line is the capped exponential delay. The dots are five clients using full jitter (each waits a random time between zero and the capped delay): instead of all retrying at the same moment, their retries spread out, which is what lets a recovering service actually recover.

In code: backoff_delays lists the capped waits $d_k$; retry_with_backoff calls a function, retries only the exceptions in RETRYABLE with those waits (optionally with full jitter), and raises anything else at once.

Brittle integrations: circuit breakers

Everyday picture. The breaker in your home's fuse box. When a circuit keeps shorting, it trips and cuts the power, instead of letting the wire overheat. After a while you flip it back on to test; if it trips again, it stays off.

A circuit breaker wraps calls to a dependency. After several consecutive failures it opens: calls fail immediately without touching the struggling service (fail fast), so the agent can fall back (use a cache, tell the user, hand off to a person) instead of hanging on timeouts. After a cool-down it goes half-open and lets one trial call through: success closes it, failure opens it again.

Worked example. Threshold 3, cool-down 30 s. Three timeouts in a row: open. A fourth call at 5 s fails instantly with CircuitOpen and the service isn't called at all. At 31 s, one trial call goes through; it succeeds, so the breaker closes.

stateDiagram-v2 [*] --> Closed Closed --> Closed: success (reset the failure count) Closed --> Open: 3 failures in a row Open --> Open: calls rejected instantly Open --> HalfOpen: cool-down elapsed HalfOpen --> Closed: trial call succeeds HalfOpen --> Open: trial call fails

Reading it: in the normal state (Closed) calls flow through and failures are counted. Open is the protective state: nothing reaches the dependency. Half-open is a single, cautious test. The breaker turns a slow, cascading failure (every request waiting 30 s on a dead service) into a fast, contained one.

Through a 50-second outage the failing service is called only 5 times, while 21 requests fail fast at the open breaker instead

Reading it: the shaded band is when the upstream service is down; each marker is one request. Before the outage, calls succeed (green). The first three failures (red) open the breaker; after that, requests fail fast (grey) without reaching the service, apart from one trial call per cool-down. Once the service is back, the next trial succeeds and traffic resumes. The service received a handful of calls during the outage instead of all of them.

In code: CircuitBreaker.call is the state diagram: it raises CircuitOpen while open, lets one trial call through after the cool-down, and closes on success. simulate_outage replays the outage in the figure.

Contract tests (automated checks that an external API still accepts and returns what your tool expects) catch the third kind of brittleness, upstream API changes, before users do.

Loops and runaway

Everyday picture. A satnav that keeps rerouting you around the same block. The fix isn't a better map; it's noticing that you've passed the same corner three times.

Worked example. search("vpn"), search("vpn"), search("vpn"): the same tool with the same arguments three times in a row. Nothing new can come back, so the agent is looping. search("vpn"), search("vpn error"), search("ERR-4012") is the same tool refining its query: progress. Combine loop detection with hard step and token budgets (primer.agents.cost.TaskBudget) so a confused agent stops and hands off instead of burning money.

In code: is_looping reports whether the last few calls were the same tool with identical arguments.

In 20 seconds

  • Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60% of the time. Shorten chains and verify at key steps.
  • Retry transient failures with capped exponential backoff and jitter; never retry invalid requests; make writes idempotent first.
  • Put circuit breakers around dependencies so outages fail fast instead of cascading.
  • Detect loops (same call, same arguments) and enforce step and token budgets.
  • Most failures are system failures (retrieval, tools, data, evals, adoption), not model failures, and each has a known fix.

Self-test questions

Your agent works in demos and fails half the time in production. Where do you look first? At the length of real tasks versus demo tasks (compounding error), and at traces for the first failing step: retrieval misses, wrong tools or arguments, tool errors from integrations, context growth. Then fix per-step reliability where it's lowest, add verification after key steps, and put the failing cases into the eval set.

Why add jitter to retries? Without it, every client that failed at the same moment retries at the same moment, producing synchronized waves of load that keep an overloaded service down. Random delays spread the retries out.

When should you not retry? When the error can't go away on its own: invalid input, permission denied, not found. And never retry a non-idempotent write without an idempotency key, or you'll duplicate its effect.

What's the difference between a retry and a circuit breaker? A retry handles a blip for one request. A breaker handles an outage across many requests: it stops sending traffic to a failing dependency, fails fast so callers can fall back, and probes for recovery.

The papers behind this lesson

No single founding paper; these are engineering patterns. The canonical write-ups are Michael Nygard's Release It! (which named the circuit breaker) and the AWS Builders' Library article on backoff and jitter, both linked below.

Further reading

on GitHub
  1r"""
  2# Why the hard ones fail: a catalogue of agent failures, and the fix for each
  3
  4Run: `python -m primer.agents.failures`
  5
  6This lesson builds on the agent loop from `primer.agents.agent_loop`, on
  7traces from `primer.agents.observability` and on the golden set from
  8`primer.agents.evals`.
  9
 10## Level 1: The practitioner's guide
 11
 12**In one sentence.** Most agent failures in production are system failures
 13with known fixes (too many steps, bad retrieval, ambiguous tools, a context
 14that rots, loops, no evals, injected instructions, messy data, brittle
 15integrations, no adoption), and this lesson is the catalogue: symptom, fix,
 16and where the fix is built.
 17
 18**When you need it.** When the agent works in demos and fails in
 19production, and nobody can say why. The tell is the gap between the two:
 20demos are three steps long and real tasks are twenty. This lesson's
 21arithmetic explains the gap on its own. If each step succeeds 95% of the
 22time, a ten-step task succeeds 60% of the time and a twenty-step task 36%;
 23at 90% per step, twenty steps succeed 12% of the time; at 99%, 82%. At 95%
 24per step, the most steps you can chain and still succeed nine times in ten
 25is two. Nothing about the model changed between the demo and production;
 26the exponent did. You need this catalogue before the first incident, as a
 27checklist, and after every incident, to name what happened. You don't need
 28the integration patterns (retries, breakers) for a tool that calls nothing
 29outside your process, and you don't need loop detection for a workflow
 30with a fixed number of steps.
 31
 32**Your options.** Six defences, from the ones that protect one request to
 33the ones that protect the whole system:
 34
 35| Option | What it does | What it guarantees | What it costs | Where it lives |
 36|---|---|---|---|---|
 37| Retries with capped exponential backoff and jitter | Wait 0.1 s, then 0.2 s, doubling to a cap, with a random spread so clients don't retry in waves | A transient blip (a timeout, a rate limit) doesn't fail the task | Latency on the retried request; a duplicate write unless the write is idempotent | Around each outside call |
 38| Circuit breaker | After a run of failures, calls fail instantly for a cool-down, then one trial call probes for recovery | An outage fails fast and stays contained instead of tying every request up in timeouts (in this lesson's 50-second outage the dead service is called 5 times and 21 requests fail fast) | A fallback path for when the breaker is open | Around each dependency |
 39| Loop detection and budgets | Stop when the same tool is called with the same arguments three times, or when a step or token budget runs out | A confused agent stops and hands off instead of burning money all night | A false stop now and then; a handoff path | The agent loop |
 40| Fewer steps, verification, checkpoints | Shorten the chain, check a step's result where it happens and retry only that step, save state to resume from the middle | Attacks the exponent directly: errors stop travelling down the chain | Design effort; a check after each key step costs a call | The task design |
 41| Contract tests | Automated checks that an outside API still accepts and returns what your tool expects | An upstream change is caught by a test, not a user | Tests to write and keep current | Your test suite |
 42| Evals and traces | A golden set in CI and a trace of every run | Regressions don't ship silently, and every failure can be replayed to its first bad step | The two lessons before this one | Your pipeline and your monitoring |
 43
 44**How to choose.** Read the symptom first; the catalogue maps each one to
 45a fix.
 46
 47- Long tasks fail and short ones don't: compounding error. Cut steps,
 48  verify after the key ones, checkpoint, and put a person at the critical
 49  points.
 50- Timeouts and expired credentials: brittle integrations. Retries with
 51  backoff for blips, a breaker for outages, contract tests for changes.
 52  Never retry an invalid request (it fails the same way every time), and
 53  never retry a write that isn't idempotent.
 54- The same call repeated, cost spiking: loops. Detect the repeat and set
 55  hard budgets.
 56- Confident wrong answers: bad retrieval, or messy source data. Fix search
 57  and ingestion before touching the prompt.
 58- Wrong tool or bad arguments: ambiguous tools. Fewer tools, precise
 59  descriptions, validated inputs.
 60- The agent obeys text it found in a document: prompt injection. Boundaries
 61  around untrusted content and privilege separation.
 62- It works and nobody uses it: adoption. Build with the people it serves.
 63- Whatever the symptom, look at the trace for the first failing step before
 64  changing anything, and add the case to the golden set.
 65
 66**What it costs.** Retries cost time on the slow path: the lesson's waits
 67are 0.1 s and 0.2 s before the third attempt succeeds, capped at 10 s
 68however many attempts there are, and with full jitter each client waits a
 69random time up to that cap. A breaker costs almost nothing while closed
 70and saves a great deal while open: every request that would have waited
 7130 s on a dead service fails at once instead. Loop detection is
 72a comparison of the last few calls. Verification after a step costs a
 73model call per check, which is why it goes after the key steps, not all of
 74them. The expensive fix is the design one, shortening chains, and it is
 75also the one with the largest payoff: raising per-step reliability from
 7695% to 99% takes a twenty-step task from 36% to 82%. The lesson's figures
 77draw each of these curves so the trade can be read off, not guessed.
 78
 79**What breaks.**
 80
 81- **Retrying the invalid.** A bad request retried five times is five times
 82  the load and the same error. Retry only the exceptions that can clear
 83  on their own.
 84- **Retrying a write without an idempotency key.** "Create order" twice is
 85  two orders.
 86- **Retries in waves.** Every client that failed at the same moment retries
 87  at the same moment, and the recovering service goes down again. Jitter.
 88- **No breaker.** One dead dependency, and every request waits out a full
 89  timeout; the whole agent looks down.
 90- **A breaker with no fallback.** Failing fast is only useful if the agent
 91  can do something else: use a cache, tell the user, hand off.
 92- **Loop detection that only counts.** Three calls of the same tool with
 93  refining queries is progress, not a loop. Compare arguments, not names.
 94- **Fixing the prompt for a system fault.** A retrieval miss, a tool error
 95  or a parsing bug looks like a model mistake from the answer alone. The
 96  trace shows which it was.
 97
 98**In the wild.** The circuit breaker was named by Michael Nygard in
 99*Release It!* and written up by Martin Fowler with the three states this
100lesson builds (closed, open, half-open); the retry guidance, capped
101exponential backoff with jitter, is the AWS Builders' Library article the
102lesson links. Libraries package both: in
103Python, tenacity gives a retry decorator with exponential and random
104exponential waits, stop conditions by attempts or elapsed time, and retry
105only on chosen exception types, the same policy as this lesson's
106`retry_with_backoff`. Anthropic's *Building effective agents* draws the
107same conclusions from the other direction: start with simple prompts and
108evaluation, add multi-step agents only where simpler designs fall short,
109add complexity only when it demonstrably improves outcomes, and have the
110agent pause for human feedback at checkpoints; its emphasis on tool design
111as an interface problem is the fix for the ambiguous-tools row. The seven
112rows this lesson doesn't build point to the lessons that do.
113
114**Go deeper.** Level 2 builds three of the fixes in plain code: the
115compounding-error formula and its inverse (how many steps a target allows),
116retries with capped backoff and full jitter that raise invalid requests
117at once, a circuit breaker as a three-state machine replayed through an
118outage, and a loop detector that compares arguments. If you only needed to
119name the failure and find its fix, you are done.
120
121## Level 2: How it works, from scratch
122
123Pilots learn from a catalogue of accidents: each entry names what went
124wrong, how it showed up in the cockpit, and the checklist item that now
125prevents it. This lesson is that catalogue for agents: ten failures that
126keep recurring in production systems, each with its symptom, its fix, and a
127pointer to the lesson that builds the fix. Three of them get their fix
128built right here: compounding error (arithmetic), brittle integrations
129(retries with backoff, and circuit breakers) and runaway loops (loop
130detection).
131
132| Failure | Symptom | Fix | Lesson |
133|---|---|---|---|
134| Compounding error | Long tasks fail far more often than short ones | Fewer steps, checks after key steps, checkpoints | `primer.agents.planning` |
135| Bad retrieval | Confident, wrong answers | Hybrid search, reranking, retrieval evals | `primer.agents.rag` |
136| Ambiguous tools | Wrong tool, or bad arguments | Precise descriptions, fewer tools, validation | `primer.agents.tools` |
137| Context rot | Quality drops as a session grows | Summarize, trim, restart with a state handoff | `primer.agents.context` |
138| Loops and runaway | The same call repeated; cost spikes | Step budgets, loop detection, stop conditions | `primer.agents.agent_loop` |
139| No evals | Regressions ship silently | Golden sets in CI, online monitoring | `primer.agents.evals` |
140| Prompt injection | The agent obeys instructions found in data | Untrusted-content boundaries, privilege separation | `primer.agents.guardrails` |
141| Messy enterprise data | Garbled tables, missing permissions | Invest in parsing; permission-aware retrieval | `primer.agents.rag` |
142| Brittle integrations | Timeouts, expired credentials, API changes | Retries with backoff, circuit breakers, contract tests | `primer.agents.failures` |
143| No adoption | It works, and nobody uses it | Build with users, show sources, easy human handoff | `primer.agents.deployment` |
144
145**In code:** `FAILURES` is this table as data, one `Failure` record per row.
146
147## Compounding error
148
149**Everyday picture.** A relay race where each baton pass succeeds 95% of
150the time. One pass is nearly safe; ten passes in a row drop the baton more
151often than you'd think.
152
153**Worked example.** If each step of an agent's task succeeds 95% of the
154time, independently:
155
156| Steps | Chance every step succeeds |
157|---|---|
158| 1 | 0.95 |
159| 2 | 0.95 × 0.95 = 0.9025 |
160| 10 | 0.95¹⁰ ≈ **0.599** |
161| 20 | 0.95²⁰ ≈ **0.358** |
162
163$$
164P(\text{task succeeds}) = p^{\,n}
165$$
166
167**Symbols**
168
169| Symbol | Meaning here | Worked example |
170|---|---|---|
171| $p$ | chance a single step succeeds | 0.95 |
172| $n$ | number of steps that must all succeed | 10 |
173| $p^{\,n}$ | $p$ multiplied by itself $n$ times | 0.599 |
174
175**In words:** the chance that every step succeeds is the per-step chance
176multiplied together once per step.
177
178**On the worked example:** 0.95 multiplied by itself 10 times is 0.599, so a
179ten-step task at 95% per step fails four times in ten.
180
181**In Python:**
182
183```python
184p, n = 0.95, 10
185# p multiplied by itself n times
186round(p ** n, 3)  # → 0.599
187# fails about four times in ten
188round(1 - p ** n, 1)  # → 0.4
189```
190
191```mermaid
192flowchart LR
193  S1[Step 1<br/>95%] --> S2[Step 2<br/>95%] --> S3[...] --> S10[Step 10<br/>95%] --> D[Done:<br/>60% of the time]
194  S2 -.->|verify, retry<br/>just this step| S2
195```
196
197**Reading it:** every arrow is another chance to drop the baton, so success
198shrinks with each step. The dotted loop is the remedy: check a step's
199result where it happens and retry only that step, so an error doesn't
200travel down the chain. Fewer steps, verification after key steps,
201checkpoints to resume from the middle, and human review at critical points
202all attack the same exponent.
203
204![Twenty steps succeed 82% of the time at 99% per step but only 36% at 95% and 12% at 90%](figures/primer.agents.failures.compounding.svg)
205
206**Reading it:** the x-axis is the number of steps; each curve is a per-step
207reliability. At 99% per step a 20-step task still succeeds 82% of the time;
208at 95% it's 36%; at 90%, 12%. Small gains in per-step reliability are worth
209far more than they look, and demos with three steps say little about tasks
210with twenty.
211
212**In code:** `chain_success` is $p^{\,n}$; `max_steps_for` runs it backwards,
213returning the most steps you can chain and still reach a target success rate.
214
215## Brittle integrations: retries with backoff
216
217**Everyday picture.** Calling a busy phone line: you don't redial every
218second. You wait a little, then longer, then longer still, and if everyone
219redials on the same schedule the line stays jammed, so you add a random
220pause (**jitter**).
221
222**Worked example.** A service times out twice and then answers. With a base
223delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt
224succeeds. With a cap of 10 s, the waits never exceed 10 s however many
225attempts there are. A request that's *invalid* (a `ValueError` here) fails
226the same way every time, so it's raised immediately: retrying only adds
227load.
228
229$$
230d_k = \min\bigl(D_{\max},\; b \cdot 2^{k}\bigr), \qquad k = 0, 1, 2, \ldots
231$$
232
233**Symbols**
234
235| Symbol | Meaning here | Worked example |
236|---|---|---|
237| $d_k$ | wait before retry number $k+1$ | $d_0$ = 0.1 s, $d_1$ = 0.2 s |
238| $k$ | how many retries have already happened | 0, 1 |
239| $b$ | base delay | 0.1 s |
240| $2^{k}$ | doubles with every retry | 1, 2, 4, … |
241| $D_{\max}$ | the cap on any single wait | 10 s |
242
243**In words:** each wait doubles the previous one, starting from the base
244delay, but never exceeds the cap.
245
246**On the worked example:** $d_0$ = min(10, 0.1 × 1) = 0.1 s and
247$d_1$ = min(10, 0.1 × 2) = 0.2 s.
248
249**In Python:**
250
251```python
252b, D_max = 0.1, 10
253# d_0 and d_1
254[min(D_max, b * 2 ** k) for k in range(2)]  # → [0.1, 0.2]
255```
256
257```mermaid
258sequenceDiagram
259  participant A as Agent tool
260  participant S as Upstream API
261  A->>S: request
262  S--xA: timeout
263  Note over A: wait 0.1 s
264  A->>S: retry 1
265  S--xA: timeout
266  Note over A: wait 0.2 s
267  A->>S: retry 2
268  S-->>A: 200 OK
269```
270
271**Reading it:** time runs downward. Each failure is followed by a longer
272pause, which gives an overloaded service room to recover instead of being
273hammered. Writes must be **idempotent** (safe to repeat: sending the same
274request twice has the same effect as once, usually via an idempotency key)
275before you retry them, or a retried "create order" makes two orders.
276
277![The delay doubles each attempt until it hits the 10-second cap, and jitter scatters five clients' retries across the whole range below it](figures/primer.agents.failures.backoff.svg)
278
279**Reading it:** the x-axis is the attempt number; the black line is the
280capped exponential delay. The dots are five clients using *full jitter*
281(each waits a random time between zero and the capped delay): instead of
282all retrying at the same moment, their retries spread out, which is what
283lets a recovering service actually recover.
284
285**In code:** `backoff_delays` lists the capped waits $d_k$;
286`retry_with_backoff` calls a function, retries only the exceptions in
287`RETRYABLE` with those waits (optionally with full jitter), and raises
288anything else at once.
289
290## Brittle integrations: circuit breakers
291
292**Everyday picture.** The breaker in your home's fuse box. When a circuit
293keeps shorting, it trips and cuts the power, instead of letting the wire
294overheat. After a while you flip it back on to test; if it trips again, it
295stays off.
296
297A **circuit breaker** wraps calls to a dependency. After several
298consecutive failures it **opens**: calls fail immediately without touching
299the struggling service (fail fast), so the agent can fall back (use a cache,
300tell the user, hand off to a person) instead of hanging on timeouts. After a
301cool-down it goes **half-open** and lets one trial call through: success
302closes it, failure opens it again.
303
304**Worked example.** Threshold 3, cool-down 30 s. Three timeouts in a row:
305open. A fourth call at 5 s fails instantly with `CircuitOpen` and the
306service isn't called at all. At 31 s, one trial call goes through; it
307succeeds, so the breaker closes.
308
309```mermaid
310stateDiagram-v2
311  [*] --> Closed
312  Closed --> Closed: success (reset the failure count)
313  Closed --> Open: 3 failures in a row
314  Open --> Open: calls rejected instantly
315  Open --> HalfOpen: cool-down elapsed
316  HalfOpen --> Closed: trial call succeeds
317  HalfOpen --> Open: trial call fails
318```
319
320**Reading it:** in the normal state (Closed) calls flow through and failures
321are counted. Open is the protective state: nothing reaches the dependency.
322Half-open is a single, cautious test. The breaker turns a slow, cascading
323failure (every request waiting 30 s on a dead service) into a fast,
324contained one.
325
326![Through a 50-second outage the failing service is called only 5 times, while 21 requests fail fast at the open breaker instead](figures/primer.agents.failures.breaker.svg)
327
328**Reading it:** the shaded band is when the upstream service is down;
329each marker is one request. Before the outage, calls succeed (green). The
330first three failures (red) open the breaker; after that, requests fail fast
331(grey) without reaching the service, apart from one trial call per
332cool-down. Once the service is back, the next trial succeeds and traffic
333resumes. The service received a handful of calls during the outage instead
334of all of them.
335
336**In code:** `CircuitBreaker.call` is the state diagram: it raises
337`CircuitOpen` while open, lets one trial call through after the cool-down,
338and closes on success. `simulate_outage` replays the outage in the figure.
339
340**Contract tests** (automated checks that an external API still accepts
341and returns what your tool expects) catch the third kind of brittleness,
342upstream API changes, before users do.
343
344## Loops and runaway
345
346**Everyday picture.** A satnav that keeps rerouting you around the same
347block. The fix isn't a better map; it's noticing that you've passed the same
348corner three times.
349
350**Worked example.** `search("vpn")`, `search("vpn")`, `search("vpn")`: the
351same tool with the same arguments three times in a row. Nothing new can
352come back, so the agent is looping. `search("vpn")`, `search("vpn error")`,
353`search("ERR-4012")` is the same tool refining its query: progress.
354Combine loop detection with hard step and token budgets
355(`primer.agents.cost.TaskBudget`) so a confused agent stops and hands off
356instead of burning money.
357
358**In code:** `is_looping` reports whether the last few calls were the same
359tool with identical arguments.
360
361## In 20 seconds
362- Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60%
363  of the time. Shorten chains and verify at key steps.
364- Retry transient failures with capped exponential backoff and jitter; never
365  retry invalid requests; make writes idempotent first.
366- Put circuit breakers around dependencies so outages fail fast instead of
367  cascading.
368- Detect loops (same call, same arguments) and enforce step and token
369  budgets.
370- Most failures are system failures (retrieval, tools, data, evals,
371  adoption), not model failures, and each has a known fix.
372
373## Self-test questions
374
375**Your agent works in demos and fails half the time in production. Where do
376you look first?**
377At the length of real tasks versus demo tasks (compounding error), and at
378traces for the first failing step: retrieval misses, wrong tools or
379arguments, tool errors from integrations, context growth. Then fix
380per-step reliability where it's lowest, add verification after key steps,
381and put the failing cases into the eval set.
382
383**Why add jitter to retries?**
384Without it, every client that failed at the same moment retries at the same
385moment, producing synchronized waves of load that keep an overloaded
386service down. Random delays spread the retries out.
387
388**When should you not retry?**
389When the error can't go away on its own: invalid input, permission denied,
390not found. And never retry a non-idempotent write without an idempotency
391key, or you'll duplicate its effect.
392
393**What's the difference between a retry and a circuit breaker?**
394A retry handles a blip for one request. A breaker handles an outage across
395many requests: it stops sending traffic to a failing dependency, fails fast
396so callers can fall back, and probes for recovery.
397
398## The papers behind this lesson
399
400No single founding paper; these are engineering patterns. The canonical
401write-ups are Michael Nygard's *Release It!* (which named the circuit
402breaker) and the AWS Builders' Library article on backoff and jitter, both
403linked below.
404
405## Further reading
406- Martin Fowler, *CircuitBreaker*: https://martinfowler.com/bliki/CircuitBreaker.html
407- AWS Builders' Library, *Timeouts, retries, and backoff with jitter*: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
408- Anthropic, *Building effective agents*: https://www.anthropic.com/engineering/building-effective-agents
409"""
410
411from __future__ import annotations
412
413import json
414import math
415import time
416from dataclasses import dataclass
417from typing import Any, Callable
418
419import numpy as np
420
421from primer._show import banner, say, table, takeaway
422
423# ---------------------------------------------------------------------------
424# 1. The catalogue
425# ---------------------------------------------------------------------------
426
427
428@dataclass(frozen=True)
429class Failure:
430    name: str
431    symptom: str
432    fix: str
433    module: str
434
435
436FAILURES: list[Failure] = [
437    Failure("Compounding error", "Long tasks fail far more than short ones", "Fewer steps, verify key steps, checkpoints", "primer.agents.planning"),
438    Failure("Bad retrieval", "Confident wrong answers", "Hybrid search, reranking, retrieval evals", "primer.agents.rag"),
439    Failure("Ambiguous tools", "Wrong tool or bad arguments", "Precise descriptions, fewer tools, validation", "primer.agents.tools"),
440    Failure("Context rot", "Quality drops as sessions grow", "Summarize, trim, restart with a state handoff", "primer.agents.context"),
441    Failure("Loops and runaway", "Repeated calls, cost spikes", "Step budgets, loop detection, stop conditions", "primer.agents.agent_loop"),
442    Failure("No evals", "Regressions ship silently", "Golden sets in CI, online monitoring", "primer.agents.evals"),
443    Failure("Prompt injection", "Agent follows instructions found in data", "Untrusted-content boundaries, privilege separation", "primer.agents.guardrails"),
444    Failure("Messy enterprise data", "Garbled parsing, permission gaps", "Invest in ingestion, permission-aware retrieval", "primer.agents.rag"),
445    Failure("Brittle integrations", "Timeouts, auth expiry, API changes", "Retries with backoff, circuit breakers, contract tests", "primer.agents.failures"),
446    Failure("No adoption", "Works, but nobody uses it", "Design with users, show sources, easy human handoff", "primer.agents.deployment"),
447]
448
449# ---------------------------------------------------------------------------
450# 2. Compounding error
451# ---------------------------------------------------------------------------
452
453
454def chain_success(p: float, n: int) -> float:
455    """Chance that n independent steps, each succeeding with probability p, all succeed."""
456    return p**n
457
458
459def max_steps_for(p: float, target: float) -> int:
460    """The most steps you can chain at per-step reliability p and still reach `target`."""
461    # p**n >= target  <=>  n <= log(target) / log(p)   (both logs are negative)
462    return math.floor(math.log(target) / math.log(p) + 1e-12)
463
464
465# ---------------------------------------------------------------------------
466# 3. Retries with exponential backoff
467# ---------------------------------------------------------------------------
468
469RETRYABLE: tuple[type[Exception], ...] = (TimeoutError, ConnectionError)
470
471
472def backoff_delays(attempts: int, base_delay: float, max_delay: float) -> list[float]:
473    """The capped exponential wait before each retry (no jitter)."""
474    return [min(max_delay, base_delay * 2**k) for k in range(attempts - 1)]
475
476
477def retry_with_backoff(fn: Callable[[], Any], max_attempts: int = 5, base_delay: float = 0.1, max_delay: float = 10.0,
478                       jitter: np.random.Generator | None = None, sleep: Callable[[float], None] = time.sleep,
479                       retryable: tuple[type[Exception], ...] = RETRYABLE) -> Any:
480    """Call fn, retrying transient failures with capped exponential backoff.
481
482    Pass a seeded `np.random.default_rng()` as `jitter` for "full jitter":
483    wait a uniform random time between 0 and the capped delay.
484    """
485    for attempt in range(max_attempts):
486        try:
487            return fn()
488        except retryable:
489            if attempt == max_attempts - 1:
490                raise
491            delay = min(max_delay, base_delay * 2**attempt)
492            sleep(float(jitter.uniform(0, delay)) if jitter is not None else delay)
493    raise AssertionError("unreachable")
494
495
496# ---------------------------------------------------------------------------
497# 4. Circuit breaker
498# ---------------------------------------------------------------------------
499
500
501class CircuitOpen(RuntimeError):
502    """Raised instead of calling a dependency whose breaker is open."""
503
504
505class CircuitBreaker:
506    def __init__(self, failure_threshold: int = 3, reset_timeout: float = 30.0, clock: Callable[[], float] = time.monotonic):
507        self.threshold, self.reset_timeout, self.clock = failure_threshold, reset_timeout, clock
508        self.state = "closed"
509        self.failures = 0
510        self.opened_at = 0.0
511
512    def call(self, fn: Callable[[], Any]) -> Any:
513        if self.state == "open":
514            if self.clock() - self.opened_at < self.reset_timeout:
515                raise CircuitOpen("dependency unavailable; failing fast")
516            self.state = "half_open"  # cool-down over: allow one trial call
517        try:
518            result = fn()
519        except Exception:
520            self.failures += 1
521            if self.state == "half_open" or self.failures >= self.threshold:
522                self.state, self.opened_at = "open", self.clock()
523            raise
524        self.state, self.failures = "closed", 0
525        return result
526
527
528def simulate_outage(outage: tuple[float, float] = (10.0, 60.0), every: float = 2.0, until: float = 100.0) -> list[dict[str, Any]]:
529    """Requests every `every` seconds through a breaker while the service is down during `outage`."""
530    now = [0.0]
531    breaker = CircuitBreaker(failure_threshold=3, reset_timeout=15.0, clock=lambda: now[0])
532    events = []
533
534    def service():
535        if outage[0] <= now[0] < outage[1]:
536            raise TimeoutError("down")
537        return "ok"
538
539    t = 0.0
540    while t < until:
541        now[0] = t
542        try:
543            breaker.call(service)
544            outcome = "ok"
545        except CircuitOpen:
546            outcome = "fast-fail"
547        except TimeoutError:
548            outcome = "failed"
549        events.append({"t": t, "outcome": outcome, "state": breaker.state})
550        t += every
551    return events
552
553
554# ---------------------------------------------------------------------------
555# 5. Loop detection
556# ---------------------------------------------------------------------------
557
558
559def is_looping(calls: list[tuple[str, dict[str, Any]]], repeats: int = 3) -> bool:
560    """True if the last `repeats` calls are the same tool with identical arguments."""
561    if len(calls) < repeats:
562        return False
563    keys = [(name, json.dumps(args, sort_keys=True)) for name, args in calls[-repeats:]]
564    return len(set(keys)) == 1
565
566
567# ---------------------------------------------------------------------------
568# Figures and demo
569# ---------------------------------------------------------------------------
570
571
572def figures() -> dict[str, Any]:
573    import matplotlib
574
575    matplotlib.use("Agg")
576    import matplotlib.pyplot as plt
577
578    figs: dict[str, Any] = {}
579
580    n = np.arange(1, 31)
581    fig, ax = plt.subplots(figsize=(6.5, 3.5))
582    for p, color in [(0.99, "#55a868"), (0.95, "#4c72b0"), (0.90, "#c44e52")]:
583        ax.plot(n, [chain_success(p, k) for k in n], color=color, label=f"{p:.0%} per step")
584    for k in (10, 20):
585        ax.plot(k, chain_success(0.95, k), "o", color="#4c72b0")
586        ax.annotate(f"{chain_success(0.95, k):.0%}", (k, chain_success(0.95, k)), textcoords="offset points", xytext=(5, 5), fontsize=8)
587    ax.set_xlabel("steps that must all succeed")
588    ax.set_ylabel("chance the whole task succeeds")
589    ax.set_ylim(0, 1.02)
590    ax.set_title("Compounding error")
591    ax.legend()
592    fig.tight_layout()
593    figs["compounding"] = fig
594
595    attempts = 8
596    base = backoff_delays(attempts, 0.5, 10.0)
597    rng = np.random.default_rng(0)
598    fig, ax = plt.subplots(figsize=(6.5, 3.5))
599    ax.plot(range(1, attempts), base, "k-o", label="capped exponential (no jitter)")
600    for c in range(5):
601        ax.scatter(np.arange(1, attempts) + (c - 2) * 0.06, [rng.uniform(0, d) for d in base], s=18, alpha=0.8,
602                   label="five clients with full jitter" if c == 0 else None, color="#4c72b0")
603    ax.set_xlabel("retry number")
604    ax.set_ylabel("seconds to wait")
605    ax.set_title("Backoff: base 0.5 s, doubling, capped at 10 s")
606    ax.legend(fontsize=8)
607    fig.tight_layout()
608    figs["backoff"] = fig
609
610    ev = simulate_outage()
611    fig, ax = plt.subplots(figsize=(7, 2.8))
612    ax.axvspan(10, 60, color="#f2d0d0", label="service down")
613    colors = {"ok": "#55a868", "failed": "#c44e52", "fast-fail": "#8c8c8c"}
614    for kind in colors:
615        pts = [e["t"] for e in ev if e["outcome"] == kind]
616        ax.scatter(pts, [1] * len(pts), color=colors[kind], s=40, label=kind.replace("fast-fail", "failed fast (breaker open)"))
617    ax.set_yticks([])
618    ax.set_xlabel("seconds")
619    ax.set_title("Circuit breaker (threshold 3, cool-down 15 s) during an outage")
620    ax.legend(fontsize=7, loc="upper center", bbox_to_anchor=(0.5, -0.35), ncol=4, frameon=False)
621    fig.tight_layout()
622    figs["breaker"] = fig
623    return figs
624
625
626def demo() -> None:
627    banner("1. The catalogue")
628    table(["failure", "symptom", "fix", "lesson"], [(f.name, f.symptom, f.fix, f.module) for f in FAILURES])
629
630    banner("2. Compounding error")
631    table(["steps", "at 99%", "at 95%", "at 90%"], [(k, chain_success(0.99, k), chain_success(0.95, k), chain_success(0.90, k)) for k in (1, 5, 10, 20)], floatfmt=".3f")
632    print(f"at 95% per step, the most steps that keep success at 90% or better: {max_steps_for(0.95, 0.90)}")
633    print()
634
635    banner("3. Retries with backoff")
636    waits: list[float] = []
637    calls = {"n": 0}
638
639    def flaky():
640        calls["n"] += 1
641        # Two timeouts, then an answer: the worked example in the lesson.
642        if calls["n"] <= 2:
643            raise TimeoutError("upstream timed out")
644        return "ok"
645
646    print("result:", retry_with_backoff(flaky, base_delay=0.1, sleep=waits.append), "after waits", waits)
647    print()
648
649    banner("4. Circuit breaker during an outage")
650    ev = simulate_outage()
651    reached = sum(e["outcome"] == "failed" for e in ev)
652    fast = sum(e["outcome"] == "fast-fail" for e in ev)
653    print(f"during the outage the service was hit {reached} times; {fast} calls failed fast instead")
654    print()
655
656    banner("5. Loop detection")
657    print(is_looping([("search", {"q": "vpn"})] * 3), is_looping([("search", {"q": "vpn"}), ("search", {"q": "vpn error"}), ("search", {"q": "ERR-4012"})]))
658    print()
659    say("Same tool, same arguments, three times: nothing new can come back. Stop, and hand off.")
660    takeaway("Most agent failures are system failures, and each one has a known fix. Build the fix, then test for it.")
661
662
663if __name__ == "__main__":
664    demo()
Level 3: the code, function by function.
@dataclass(frozen=True)
class Failure: on GitHub
429@dataclass(frozen=True)
430class Failure:
431    name: str
432    symptom: str
433    fix: str
434    module: str
Failure(name: str, symptom: str, fix: str, module: str)
name: str
symptom: str
fix: str
module: str
FAILURES: list[Failure] = [Failure(name='Compounding error', symptom='Long tasks fail far more than short ones', fix='Fewer steps, verify key steps, checkpoints', module='primer.agents.planning'), Failure(name='Bad retrieval', symptom='Confident wrong answers', fix='Hybrid search, reranking, retrieval evals', module='primer.agents.rag'), Failure(name='Ambiguous tools', symptom='Wrong tool or bad arguments', fix='Precise descriptions, fewer tools, validation', module='primer.agents.tools'), Failure(name='Context rot', symptom='Quality drops as sessions grow', fix='Summarize, trim, restart with a state handoff', module='primer.agents.context'), Failure(name='Loops and runaway', symptom='Repeated calls, cost spikes', fix='Step budgets, loop detection, stop conditions', module='primer.agents.agent_loop'), Failure(name='No evals', symptom='Regressions ship silently', fix='Golden sets in CI, online monitoring', module='primer.agents.evals'), Failure(name='Prompt injection', symptom='Agent follows instructions found in data', fix='Untrusted-content boundaries, privilege separation', module='primer.agents.guardrails'), Failure(name='Messy enterprise data', symptom='Garbled parsing, permission gaps', fix='Invest in ingestion, permission-aware retrieval', module='primer.agents.rag'), Failure(name='Brittle integrations', symptom='Timeouts, auth expiry, API changes', fix='Retries with backoff, circuit breakers, contract tests', module='primer.agents.failures'), Failure(name='No adoption', symptom='Works, but nobody uses it', fix='Design with users, show sources, easy human handoff', module='primer.agents.deployment')]
def chain_success(p: float, n: int) -> float: on GitHub
455def chain_success(p: float, n: int) -> float:
456    """Chance that n independent steps, each succeeding with probability p, all succeed."""
457    return p**n

Chance that n independent steps, each succeeding with probability p, all succeed.

def max_steps_for(p: float, target: float) -> int: on GitHub
460def max_steps_for(p: float, target: float) -> int:
461    """The most steps you can chain at per-step reliability p and still reach `target`."""
462    # p**n >= target  <=>  n <= log(target) / log(p)   (both logs are negative)
463    return math.floor(math.log(target) / math.log(p) + 1e-12)

The most steps you can chain at per-step reliability p and still reach target.

RETRYABLE: tuple[type[Exception], ...] = (<class 'TimeoutError'>, <class 'ConnectionError'>)
def backoff_delays(attempts: int, base_delay: float, max_delay: float) -> list[float]: on GitHub
473def backoff_delays(attempts: int, base_delay: float, max_delay: float) -> list[float]:
474    """The capped exponential wait before each retry (no jitter)."""
475    return [min(max_delay, base_delay * 2**k) for k in range(attempts - 1)]

The capped exponential wait before each retry (no jitter).

def retry_with_backoff( fn: Callable[[], Any], max_attempts: int = 5, base_delay: float = 0.1, max_delay: float = 10.0, jitter: numpy.random._generator.Generator | None = None, sleep: Callable[[float], NoneType] = <built-in function sleep>, retryable: tuple[type[Exception], ...] = (<class 'TimeoutError'>, <class 'ConnectionError'>)) -> Any: on GitHub
478def retry_with_backoff(fn: Callable[[], Any], max_attempts: int = 5, base_delay: float = 0.1, max_delay: float = 10.0,
479                       jitter: np.random.Generator | None = None, sleep: Callable[[float], None] = time.sleep,
480                       retryable: tuple[type[Exception], ...] = RETRYABLE) -> Any:
481    """Call fn, retrying transient failures with capped exponential backoff.
482
483    Pass a seeded `np.random.default_rng()` as `jitter` for "full jitter":
484    wait a uniform random time between 0 and the capped delay.
485    """
486    for attempt in range(max_attempts):
487        try:
488            return fn()
489        except retryable:
490            if attempt == max_attempts - 1:
491                raise
492            delay = min(max_delay, base_delay * 2**attempt)
493            sleep(float(jitter.uniform(0, delay)) if jitter is not None else delay)
494    raise AssertionError("unreachable")

Call fn, retrying transient failures with capped exponential backoff.

Pass a seeded np.random.default_rng() as jitter for "full jitter": wait a uniform random time between 0 and the capped delay.

class CircuitOpen(builtins.RuntimeError): on GitHub
502class CircuitOpen(RuntimeError):
503    """Raised instead of calling a dependency whose breaker is open."""

Raised instead of calling a dependency whose breaker is open.

class CircuitBreaker: on GitHub
506class CircuitBreaker:
507    def __init__(self, failure_threshold: int = 3, reset_timeout: float = 30.0, clock: Callable[[], float] = time.monotonic):
508        self.threshold, self.reset_timeout, self.clock = failure_threshold, reset_timeout, clock
509        self.state = "closed"
510        self.failures = 0
511        self.opened_at = 0.0
512
513    def call(self, fn: Callable[[], Any]) -> Any:
514        if self.state == "open":
515            if self.clock() - self.opened_at < self.reset_timeout:
516                raise CircuitOpen("dependency unavailable; failing fast")
517            self.state = "half_open"  # cool-down over: allow one trial call
518        try:
519            result = fn()
520        except Exception:
521            self.failures += 1
522            if self.state == "half_open" or self.failures >= self.threshold:
523                self.state, self.opened_at = "open", self.clock()
524            raise
525        self.state, self.failures = "closed", 0
526        return result
CircuitBreaker( failure_threshold: int = 3, reset_timeout: float = 30.0, clock: Callable[[], float] = <built-in function monotonic>) on GitHub
507    def __init__(self, failure_threshold: int = 3, reset_timeout: float = 30.0, clock: Callable[[], float] = time.monotonic):
508        self.threshold, self.reset_timeout, self.clock = failure_threshold, reset_timeout, clock
509        self.state = "closed"
510        self.failures = 0
511        self.opened_at = 0.0
state
failures
def call(self, fn: Callable[[], Any]) -> Any: on GitHub
513    def call(self, fn: Callable[[], Any]) -> Any:
514        if self.state == "open":
515            if self.clock() - self.opened_at < self.reset_timeout:
516                raise CircuitOpen("dependency unavailable; failing fast")
517            self.state = "half_open"  # cool-down over: allow one trial call
518        try:
519            result = fn()
520        except Exception:
521            self.failures += 1
522            if self.state == "half_open" or self.failures >= self.threshold:
523                self.state, self.opened_at = "open", self.clock()
524            raise
525        self.state, self.failures = "closed", 0
526        return result
def simulate_outage( outage: tuple[float, float] = (10.0, 60.0), every: float = 2.0, until: float = 100.0) -> list[dict[str, typing.Any]]: on GitHub
529def simulate_outage(outage: tuple[float, float] = (10.0, 60.0), every: float = 2.0, until: float = 100.0) -> list[dict[str, Any]]:
530    """Requests every `every` seconds through a breaker while the service is down during `outage`."""
531    now = [0.0]
532    breaker = CircuitBreaker(failure_threshold=3, reset_timeout=15.0, clock=lambda: now[0])
533    events = []
534
535    def service():
536        if outage[0] <= now[0] < outage[1]:
537            raise TimeoutError("down")
538        return "ok"
539
540    t = 0.0
541    while t < until:
542        now[0] = t
543        try:
544            breaker.call(service)
545            outcome = "ok"
546        except CircuitOpen:
547            outcome = "fast-fail"
548        except TimeoutError:
549            outcome = "failed"
550        events.append({"t": t, "outcome": outcome, "state": breaker.state})
551        t += every
552    return events

Requests every every seconds through a breaker while the service is down during outage.

def is_looping(calls: list[tuple[str, dict[str, typing.Any]]], repeats: int = 3) -> bool: on GitHub
560def is_looping(calls: list[tuple[str, dict[str, Any]]], repeats: int = 3) -> bool:
561    """True if the last `repeats` calls are the same tool with identical arguments."""
562    if len(calls) < repeats:
563        return False
564    keys = [(name, json.dumps(args, sort_keys=True)) for name, args in calls[-repeats:]]
565    return len(set(keys)) == 1

True if the last repeats calls are the same tool with identical arguments.

def figures() -> dict[str, typing.Any]: on GitHub
573def figures() -> dict[str, Any]:
574    import matplotlib
575
576    matplotlib.use("Agg")
577    import matplotlib.pyplot as plt
578
579    figs: dict[str, Any] = {}
580
581    n = np.arange(1, 31)
582    fig, ax = plt.subplots(figsize=(6.5, 3.5))
583    for p, color in [(0.99, "#55a868"), (0.95, "#4c72b0"), (0.90, "#c44e52")]:
584        ax.plot(n, [chain_success(p, k) for k in n], color=color, label=f"{p:.0%} per step")
585    for k in (10, 20):
586        ax.plot(k, chain_success(0.95, k), "o", color="#4c72b0")
587        ax.annotate(f"{chain_success(0.95, k):.0%}", (k, chain_success(0.95, k)), textcoords="offset points", xytext=(5, 5), fontsize=8)
588    ax.set_xlabel("steps that must all succeed")
589    ax.set_ylabel("chance the whole task succeeds")
590    ax.set_ylim(0, 1.02)
591    ax.set_title("Compounding error")
592    ax.legend()
593    fig.tight_layout()
594    figs["compounding"] = fig
595
596    attempts = 8
597    base = backoff_delays(attempts, 0.5, 10.0)
598    rng = np.random.default_rng(0)
599    fig, ax = plt.subplots(figsize=(6.5, 3.5))
600    ax.plot(range(1, attempts), base, "k-o", label="capped exponential (no jitter)")
601    for c in range(5):
602        ax.scatter(np.arange(1, attempts) + (c - 2) * 0.06, [rng.uniform(0, d) for d in base], s=18, alpha=0.8,
603                   label="five clients with full jitter" if c == 0 else None, color="#4c72b0")
604    ax.set_xlabel("retry number")
605    ax.set_ylabel("seconds to wait")
606    ax.set_title("Backoff: base 0.5 s, doubling, capped at 10 s")
607    ax.legend(fontsize=8)
608    fig.tight_layout()
609    figs["backoff"] = fig
610
611    ev = simulate_outage()
612    fig, ax = plt.subplots(figsize=(7, 2.8))
613    ax.axvspan(10, 60, color="#f2d0d0", label="service down")
614    colors = {"ok": "#55a868", "failed": "#c44e52", "fast-fail": "#8c8c8c"}
615    for kind in colors:
616        pts = [e["t"] for e in ev if e["outcome"] == kind]
617        ax.scatter(pts, [1] * len(pts), color=colors[kind], s=40, label=kind.replace("fast-fail", "failed fast (breaker open)"))
618    ax.set_yticks([])
619    ax.set_xlabel("seconds")
620    ax.set_title("Circuit breaker (threshold 3, cool-down 15 s) during an outage")
621    ax.legend(fontsize=7, loc="upper center", bbox_to_anchor=(0.5, -0.35), ncol=4, frameon=False)
622    fig.tight_layout()
623    figs["breaker"] = fig
624    return figs
def demo() -> None: on GitHub
627def demo() -> None:
628    banner("1. The catalogue")
629    table(["failure", "symptom", "fix", "lesson"], [(f.name, f.symptom, f.fix, f.module) for f in FAILURES])
630
631    banner("2. Compounding error")
632    table(["steps", "at 99%", "at 95%", "at 90%"], [(k, chain_success(0.99, k), chain_success(0.95, k), chain_success(0.90, k)) for k in (1, 5, 10, 20)], floatfmt=".3f")
633    print(f"at 95% per step, the most steps that keep success at 90% or better: {max_steps_for(0.95, 0.90)}")
634    print()
635
636    banner("3. Retries with backoff")
637    waits: list[float] = []
638    calls = {"n": 0}
639
640    def flaky():
641        calls["n"] += 1
642        # Two timeouts, then an answer: the worked example in the lesson.
643        if calls["n"] <= 2:
644            raise TimeoutError("upstream timed out")
645        return "ok"
646
647    print("result:", retry_with_backoff(flaky, base_delay=0.1, sleep=waits.append), "after waits", waits)
648    print()
649
650    banner("4. Circuit breaker during an outage")
651    ev = simulate_outage()
652    reached = sum(e["outcome"] == "failed" for e in ev)
653    fast = sum(e["outcome"] == "fast-fail" for e in ev)
654    print(f"during the outage the service was hit {reached} times; {fast} calls failed fast instead")
655    print()
656
657    banner("5. Loop detection")
658    print(is_looping([("search", {"q": "vpn"})] * 3), is_looping([("search", {"q": "vpn"}), ("search", {"q": "vpn error"}), ("search", {"q": "ERR-4012"})]))
659    print()
660    say("Same tool, same arguments, three times: nothing new can come back. Stop, and hand off.")
661    takeaway("Most agent failures are system failures, and each one has a known fix. Build the fix, then test for it.")