primer.agents.planning

Multi-step planning: plans, checks, retries and why long tasks fail

Run: python -m primer.agents.planning

This lesson builds on the step-at-a-time loop from primer.agents.agent_loop, on tools from primer.agents.tools, and on task state kept outside the model from primer.agents.memory.

Level 1: The practitioner's guide

In one sentence. Multi-step planning is how an agent gets a twenty-step job done when each step is only mostly reliable: write the steps down, check each one with something outside the model, save each result so a failure costs one step, and revise the plan from the point of failure rather than starting over.

When you need it. The moment a task takes more than a handful of model calls in a row. The arithmetic is unforgiving: a step that succeeds 95% of the time gives a ten-step task a 60% chance of finishing and a twenty-step task 36%; at fifty steps it is under 8%. Even a 99% step finishes a fifty-step task only 60% of the time. That is why a demo of a three-step task looks great and the same agent on a real twenty-step job disappoints. You don't need any of this for a task that is one or two calls long, or for a fixed sequence your code can run without asking a model what to do next. The tell: an agent that finishes the small cases and fails the long ones, and a bill that shows it restarting from step one after every stumble.

Your options. From the simplest to the most robust; each later row usually keeps the earlier ones:

Option What it does What it guarantees What it costs Where it lives
One step at a time The model decides the next action after each result Adapts to anything it sees Drifts on long tasks; nobody sees the intent before it acts The agent loop
A fixed workflow in code Your code chains the calls in a known order (chaining, routing, parallel branches) Predictable path, easy to test Only for tasks whose steps you know in advance Your code
Plan-and-execute with replanning The model writes the whole plan first, executes it step by step, and rewrites the rest on a failure while keeping finished work A plan a person can read before it runs; recovery that routes around the failure An extra planning call per replan, and a cap on replans Your loop around the model
Decomposition with checks Each subtask has a definition of done that code can check A failure is caught where it happens Writing a check per subtask Your code
Checkpoints Every subtask's output is saved; a retry reruns only the failed step A failure costs one step, not the run Storage for each output Your code, or a workflow engine
External verification with retries Tests, a schema validator, a database query, or a separate grader with a rubric, fed back to the model Concrete failures the model can fix; retries that actually help The check itself, and one more call per retry Outside the model
Human review at critical points A person approves before an expensive or irreversible step Mistakes stop before they cost Latency and attention Your product

How to choose. Start with the simplest thing and add a row only when the numbers say so.

  • Steps you can list in advance: a workflow in code. Anthropic's guidance is to find the simplest solution and add complexity only when needed, reserving agents for open-ended problems where the steps cannot be predicted.
  • Steps the model must discover: plan-and-execute, with replanning and a replan limit so a planner that cannot adapt doesn't loop forever. In this lesson the first plan uses a retired API, the second routes around it, and the step that already succeeded is not run again.
  • Any task longer than a few steps: decompose into subtasks with checks, and checkpoint. The expected work to finish a ten-step task at 95% per step is 10.5 step runs with checkpoints against 13.4 restarting from scratch; at fifty steps it is 53 against about 240.
  • Wherever a wrong step can hide: verify outside the model. Asked to review its own leap-year function, this lesson's model approves the bug; a test against 1900 catches it, and the second draft passes.
  • Wherever a mistake is expensive or irreversible: a person in the loop.
  • Whatever you pick, invest in the check before the retry. With a check that catches every failure and one retry, a 95% step becomes 99.75% and ten steps succeed 97.5% of the time; with a check that catches half, 77%; with no check, 60%. Retries are only as good as the check that triggers them.

What it costs. Planning costs one extra model call up front and one per replan. Checks cost engineering: a test suite, a schema, a query against the source of truth, sometimes a second model with a rubric. Retries cost a call each, but only for failures that were noticed. Checkpoints cost storage and a small amount of bookkeeping, and they are what stop the cost of a long task from growing exponentially: without them every extra step multiplies the expected work by about 1.05, which is a straight line on a log scale. Human review costs latency. Against all of this, the cost of doing nothing is a task that is abandoned and restarted, paying for the same steps over and over.

What breaks.

  • A plan reality has invalidated. The API changed, the file moved, the assumption was wrong. Treat the plan as a hypothesis and the first failed step as evidence; replan from there, keeping what is done.
  • Self-review that approves the bug. The reviewer shares the author's blind spots and can confidently reinforce a wrong answer. Put the check outside the model.
  • Retrying a failure nobody detected. A silent failure gets no retry and poisons every later step. Measure your check's detection rate; it matters more than the retry count.
  • Restarting from step one. Without checkpoints a late failure throws away all the work, and long runs become both unreliable and expensive. Save every output; rerun only the failed step.
  • A planner that loops. Cap the replans and the total steps, and report the failure instead of trying forever.
  • Subtasks that cannot be checked. "Reconcile the invoices" has no definition of done; "each invoice appears exactly once in the match" does. Make every subtask small, concrete and verifiable.

In the wild. Wang et al. (2023), Plan-and-Solve Prompting, showed that asking a model to devise a plan and then carry it out cuts the missing-step errors of reasoning straight through. Shinn et al. (2023), Reflexion, turn external feedback such as failed tests into written lessons kept in memory for the next attempt, and report 91% pass@1 on HumanEval against 80% for the base model; the gains come from signals outside the model. Anthropic's Building effective agents names the workflow patterns in the table (prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer) and insists on stopping conditions such as a maximum number of iterations. Temporal is durable execution as a product: it persists a complete event history, and when a process dies another rebuilds the state and resumes where it stopped, with local variables intact. LangGraph's persistence and checkpoints are the same idea inside an agent framework.

Go deeper. Level 2 builds the arithmetic and the remedies in plain Python: the compounding formula and its curves, a plan-and-execute loop that replans from a failed step, five subtasks with checks and the retry that reruns only one of them, the expected-cost formulas with and without checkpoints, and a reflection-versus-verification experiment on a leap-year function. If you only needed to choose, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

What follows measures the arithmetic, then builds the three remedies one mechanism at a time.

An agent that does one thing is easy. An agent that does twenty things in a row mostly fails, for a reason that is pure arithmetic. This lesson measures that arithmetic, then builds the three remedies: decompose the task into steps you can check, verify each step with something outside the model, and checkpoint so a failure costs one step instead of the whole run.

1. Compounding error: why long tasks fail

Everyday picture. A relay race with ten runners. Each runner drops the baton only 1 time in 20 during their leg. That sounds safe, but the team needs all ten legs to go cleanly, and it loses about 4 races in 10.

Tiny worked example. Each step of an agent succeeds 95% of the time.

steps chance all succeed
1 0.95
2 0.95 × 0.95 = 0.9025
10 0.95¹⁰ ≈ 0.599
20 0.95²⁰ ≈ 0.358
Level 3: the formula and its symbols

$$ P(\text{task succeeds}) = p^{\,n} $$

Symbols

Symbol Meaning
$p$ chance a single step succeeds
$n$ number of steps the task needs

In words: multiply the per-step success rate by itself once per step.

On the example: $0.95^{10} \approx 0.60$ and $0.95^{20} \approx 0.36$. A 95%-reliable step is a 60%-reliable 10-step task.

Level 3: in Python

In Python:

p = 0.95
# ten steps that must all succeed
round(p ** 10, 2)  # → 0.6
round(p ** 20, 2)  # → 0.36

At 95% per step a 10-step task succeeds only about 60% of the time, and every curve keeps sliding toward zero as steps are added

Reading it: the x-axis is the number of steps in the task, and the y-axis is the chance the whole task succeeds. Every curve starts near the top and slides down, and the lower the per-step rate, the faster it falls. Find 10 steps on the x-axis and read up to the 95% curve: about 0.6. That's why a demo of a three-step task looks great and the same agent on a real twenty-step job disappoints.

In code: end_to_end_success multiplies the per-step rate by itself once per step, which is the whole formula above.

Why it matters. The remedies all attack $p$ or $n$: fewer steps (higher-level tools, see primer.agents.tools), checks that catch failures, retries of just the failed step, and human review at critical points.

2. Plan-and-execute, with replanning

Everyday picture. Cooking from a recipe. You read the whole recipe first (the plan), then do the steps. Halfway through you discover you're out of eggs. You don't start over, and you don't pretend you have eggs. You revise the rest of the recipe with what you've got, keeping everything you've already chopped.

Tiny worked example. Goal "Reconcile Q3 invoices". The planner's first plan uses an old API. Here's the actual run:

planner  -> {"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}
execute  fetch_payments      ok
execute  fetch_invoices_v1   failed: invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2
planner  <- "Completed so far: fetch_payments. Step fetch_invoices_v1 failed: ...retired..."
planner  -> {"steps": ["fetch_invoices_v2", "match", "draft_summary"]}
execute  fetch_invoices_v2   ok
execute  match               ok
execute  draft_summary       ok     (fetch_payments was NOT run again)
flowchart TD G[Goal] --> P[Planner writes a plan] P --> X[Execute next step] X --> C{Step result<br/>meets expectations?} C -->|yes| M{More steps?} M -->|yes| X M -->|no| D[Done] C -->|no| R{Replans left?} R -->|yes| F[Tell the planner:<br/>what's done + why it failed] F --> P R -->|no| S[Stop: report the failure]

Reading it: the inner loop (execute → check → next) is the plan being followed. The outer loop back to Planner only happens on a failed check, and it carries two facts: what's already done (so work is kept) and why the step failed (so the new plan can route around it). The Replans left? diamond stops a planner that can't adapt from looping forever.

The code. plan_and_execute(planner, goal) asks for a JSON plan, runs each action, records outputs in state, and on StepFailed asks for a revised plan. Steps already in state are skipped.

In code: PlanRun is what plan_and_execute returns: every plan the planner wrote, each executed step with "ok" or its failure, and the replan count.

Why it matters. Deciding one step at a time (a pure ReAct loop, see primer.agents.agent_loop) drifts on long tasks. An upfront plan keeps the agent on track and lets a person see its intent before it acts. The risk is following a plan that reality has invalidated, which is why replanning exists.

3. Decomposition into verifiable subtasks

Everyday picture. Moving house. "Move house" isn't something you can check off, but "every box labelled", "van booked for Saturday" and "keys handed back" are. Each has a clear definition of done.

Tiny worked example. "Reconcile Q3 invoices" becomes five subtasks, each with a check:

subtask output check (definition of done)
fetch_invoices 4 invoices at least one invoice returned
fetch_payments 3 payments at least one payment returned
match 4 rows: invoiced vs paid each invoice appears exactly once
list_mismatches INV-102 (1200 vs 1100), INV-104 (450 vs 0) none
draft_summary "Q3 reconciliation: 4 invoices, 3 payments, 2 mismatches (INV-102 short by 100, INV-104 unpaid)." mentions the mismatches

When the payments API times out once, only fetch_payments runs again: attempts are {fetch_invoices: 1, fetch_payments: 2, match: 1, ...}.

flowchart LR A[fetch_invoices] --> C[match] B[fetch_payments] --> C C --> D[list_mismatches] D --> E[draft_summary] B -. timeout .-> B

Reading it: arrows show which outputs feed which step. The dotted self-loop on fetch_payments is a retry. Because every step's output is kept (a checkpoint), the retry doesn't redo fetch_invoices.

In code: Subtask pairs a step's work with its check (the definition of done); run_subtasks runs them in order, retries only the step whose check failed, and returns a RunReport of outputs and attempts. reconcile_q3 builds the five subtasks in the table.

Checkpoints: how much work a failure costs

Without checkpoints, a failure anywhere means starting again from step one. With them, you retry just the failed step. The expected number of step executions to finish an $n$-step task:

Level 3: the formula and its symbols

$$ \text{with checkpoints: } \frac{n}{p} \qquad \text{restart from scratch: } \frac{1 - p^{n}}{(1-p)\,p^{n}} $$

Symbols

Symbol Meaning
$p$ chance a single step attempt succeeds (failures are noticed)
$n$ number of steps

In words: with checkpoints each step needs on average $1/p$ attempts, so $n$ steps need $n/p$. Restarting from scratch needs $n$ successes in a row, and the expected wait for a run of $n$ successes grows roughly like $1/p^n$.

On the example: $p = 0.95$, $n = 10$: with checkpoints $10 / 0.95 \approx 10.5$ step runs; restarting from scratch $(1 - 0.599) / (0.05 \times 0.599) \approx 13.4$. At $n = 50$ it's about 52.6 vs. 240.

Level 3: in Python

In Python:

def with_checkpoints(n, p):
    # each step needs 1/p attempts on average
    return n / p
def restart_from_scratch(n, p):
    return (1 - p ** n) / ((1 - p) * p ** n)
round(with_checkpoints(10, 0.95), 1), round(restart_from_scratch(10, 0.95), 1)  # → (10.5, 13.4)
round(with_checkpoints(50, 0.95), 1), round(restart_from_scratch(50, 0.95))  # → (52.6, 240)

On the log scale, restarting from scratch climbs as a straight line (exponential growth) while checkpoints stay close to one run per step

Reading it: the x-axis is task length, and the y-axis (log scale) is how many step executions you should expect to pay for. On a log scale each gridline is ten times the one below, so exponential growth draws a straight line. The restart line is that straight line: every extra step multiplies its cost by about $1/0.95 \approx 1.05$. The checkpoint line ($n/p$, just proportional to $n$) bends over and flattens, because on a log scale going from 10 to 20 steps rises no more than going from 5 to 10. The widening gap between them is the cost of having no checkpoints, and at 50 steps it's already about 240 step runs against 53. Long tasks without checkpoints aren't just unreliable; they're expensive. Durable execution is the engineering name for this: a workflow engine that saves each step's result so a crashed run resumes where it stopped.

In code: expected_step_runs evaluates both formulas: $n/p$ with checkpoints, the restart formula without.

4. Reflection vs. external verification

Everyday picture. Proofreading your own essay versus having someone run the numbers in it. You read what you meant to write, so your own blind spots stay blind. A calculator doesn't share them.

Tiny worked example. The model writes a leap-year function:

def is_leap(year):
    return year % 4 == 0        # forgets that 1900 was not a leap year

Reflection (asking a model to review it) replies "Looks correct: years divisible by 4 are leap years." External verification (running it against known answers) replies is_leap(1900) returned True, expected False. Fed that failure, the model's second draft passes every case.

flowchart LR G[Model writes code] --> R1[Reflection:<br/>model reads its own code] R1 -->|same blind spot| A1[Approved, still wrong] G --> V[Verification:<br/>run tests / check schema / query the DB] V -->|1900 fails| F[Failure fed back] F --> G2[Second draft] G2 --> V2[Checks pass]

Reading it: two paths leave the same first draft. The top path stays inside the model and ends at "approved, still wrong". The bottom path goes through something outside the model (tests, a schema validator, a database query), and that something produces a concrete, checkable failure the model can fix.

In code: self_review is the top path (a model reads the code and approves or not); run_checks is the bottom path (it runs the code against known answers); generate_until_checks_pass loops the bottom path, feeding each failure back to the model until the checks pass.

With verification and retries, the per-step success rate rises:

Level 3: the formula and its symbols

$$ p_{\text{step}} = p \sum_{i=0}^{r} \big((1-p)\,d\big)^{i} $$

Symbols

Symbol Meaning
$p$ chance one attempt succeeds
$d$ chance the check detects a failed attempt
$r$ retries allowed after a detected failure

In words: you succeed on the first try, or fail and notice and succeed on the next try, and so on. Failures you don't notice get no retry.

On the example: $p = 0.95$, perfect check $d = 1$, one retry: $0.95 + 0.05 \times 0.95 = 0.9975$ per step, so ten steps succeed $0.9975^{10} \approx 97.5\%$ of the time instead of 60%. With a check that catches only half the failures ($d = 0.5$): $0.97375^{10} \approx 77\%$.

Level 3: in Python

In Python:

def p_step(p, d, r):
    # p Σ_{i=0}^{r} ((1 - p) d)^i
    return p * sum(((1 - p) * d) ** i for i in range(r + 1))
# a perfect check, one retry
round(p_step(0.95, 1, 1), 4)  # → 0.9975
# ten such steps
round(p_step(0.95, 1, 1) ** 10, 3)  # → 0.975
# a check that catches half the failures
round(p_step(0.95, 0.5, 1), 5)  # → 0.97375
round(p_step(0.95, 0.5, 1) ** 10, 2)  # → 0.77

With the same 95% step, a ten-step task succeeds 97.5% of the time when a check catches every failure and allows one retry, 77% when the check catches half, and 60% with no checks

Reading it: all curves use the same 95%-reliable step. The bottom curve has no checks. The middle ones add a retry after a failure is caught, with a check that catches half or all failures. The top curve allows two retries. The gap between "half" and "all" is the lesson: retries are only as good as the check that triggers them. Invest in the check.

In code: step_success computes $p_{\text{step}}$ from $p$, $d$ and $r$; feed its result into end_to_end_success to get the ten-step curves.

In 20 seconds

  • Success compounds: $p^n$. 95% per step is 60% over ten steps.
  • Plan first, execute step by step, replan from the point of failure, and keep finished work.
  • Decompose into subtasks with explicit checks; retry only the failed step (checkpoints).
  • Self-review shares the model's blind spots; tests, schemas and database checks don't.
  • Retries help only for failures you detect, so the check matters more than the retry.

Self-test questions

Q: Each step of your agent is 95% reliable. How reliable is a 20-step task, and what do you do about it? A: $0.95^{20} \approx 36\%$. Cut the steps (higher-level tools), add external checks after key steps with a retry of just that step, checkpoint so failures don't restart the run, and put human review where a mistake is expensive. With perfect checks and one retry, per-step reliability becomes 99.75%, and 20 steps succeed about 95% of the time.

Q: Why isn't "ask the model to double-check its work" enough? A: The reviewer shares the author's blind spots, and self-critique can confidently reinforce a wrong answer. Checks outside the model (tests, schema validation, a query against the source of truth, a separate grader model with a rubric) produce concrete failures the model can act on.

Q: When is plan-and-execute better than deciding one step at a time? A: For long, multi-part tasks where staying on track matters and where showing the plan to a person before acting is valuable. Always pair it with replanning. A plan is a hypothesis, and the first failed step is evidence.

Q: What makes a subtask "good"? A: It's small, concrete and verifiable. It has a definition of done that code can check, and its output is saved so a later failure doesn't redo it.

The papers behind this lesson

  • Wang et al., Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models (2023). https://arxiv.org/abs/2305.04091. It showed that asking a model to first devise a plan and then carry it out step by step reduces missed steps compared with reasoning straight through.
  • Shinn et al., Reflexion: Language Agents with Verbal Reinforcement Learning (2023). https://arxiv.org/abs/2303.11366. Agents that turn feedback from the environment (failed tests, wrong answers) into written lessons and retry improve markedly, and the gains come from external signals. annotated companion

Further reading

on GitHub
  1r"""
  2# Multi-step planning: plans, checks, retries and why long tasks fail
  3
  4Run: `python -m primer.agents.planning`
  5
  6This lesson builds on the step-at-a-time loop from
  7`primer.agents.agent_loop`, on tools from `primer.agents.tools`, and on
  8task state kept outside the model from `primer.agents.memory`.
  9
 10## Level 1: The practitioner's guide
 11
 12**In one sentence.** Multi-step planning is how an agent gets a twenty-step
 13job done when each step is only mostly reliable: write the steps down,
 14check each one with something outside the model, save each result so a
 15failure costs one step, and revise the plan from the point of failure
 16rather than starting over.
 17
 18**When you need it.** The moment a task takes more than a handful of
 19model calls in a row. The arithmetic is unforgiving: a step that succeeds
 2095% of the time gives a ten-step task a 60% chance of finishing and a
 21twenty-step task 36%; at fifty steps it is under 8%. Even a 99% step
 22finishes a fifty-step task only 60% of the time. That is why a demo of a
 23three-step task looks great and the same agent on a real twenty-step job
 24disappoints. You don't need any of this for a task that is one or two calls
 25long, or for a fixed sequence your code can run without asking a model
 26what to do next. The tell: an agent that finishes the small cases and
 27fails the long ones, and a bill that shows it restarting from step one
 28after every stumble.
 29
 30**Your options.** From the simplest to the most robust; each later row
 31usually keeps the earlier ones:
 32
 33| Option | What it does | What it guarantees | What it costs | Where it lives |
 34|---|---|---|---|---|
 35| One step at a time | The model decides the next action after each result | Adapts to anything it sees | Drifts on long tasks; nobody sees the intent before it acts | The agent loop |
 36| A fixed workflow in code | Your code chains the calls in a known order (chaining, routing, parallel branches) | Predictable path, easy to test | Only for tasks whose steps you know in advance | Your code |
 37| Plan-and-execute with replanning | The model writes the whole plan first, executes it step by step, and rewrites the rest on a failure while keeping finished work | A plan a person can read before it runs; recovery that routes around the failure | An extra planning call per replan, and a cap on replans | Your loop around the model |
 38| Decomposition with checks | Each subtask has a definition of done that code can check | A failure is caught where it happens | Writing a check per subtask | Your code |
 39| Checkpoints | Every subtask's output is saved; a retry reruns only the failed step | A failure costs one step, not the run | Storage for each output | Your code, or a workflow engine |
 40| External verification with retries | Tests, a schema validator, a database query, or a separate grader with a rubric, fed back to the model | Concrete failures the model can fix; retries that actually help | The check itself, and one more call per retry | Outside the model |
 41| Human review at critical points | A person approves before an expensive or irreversible step | Mistakes stop before they cost | Latency and attention | Your product |
 42
 43**How to choose.** Start with the simplest thing and add a row only when
 44the numbers say so.
 45
 46- Steps you can list in advance: a workflow in code. Anthropic's guidance
 47  is to find the simplest solution and add complexity only when needed,
 48  reserving agents for open-ended problems where the steps cannot be
 49  predicted.
 50- Steps the model must discover: plan-and-execute, with replanning and a
 51  replan limit so a planner that cannot adapt doesn't loop forever. In this
 52  lesson the first plan uses a retired API, the second routes around it,
 53  and the step that already succeeded is not run again.
 54- Any task longer than a few steps: decompose into subtasks with checks,
 55  and checkpoint. The expected work to finish a ten-step task at 95% per
 56  step is 10.5 step runs with checkpoints against 13.4 restarting from
 57  scratch; at fifty steps it is 53 against about 240.
 58- Wherever a wrong step can hide: verify outside the model. Asked to
 59  review its own leap-year function, this lesson's model approves the bug;
 60  a test against 1900 catches it, and the second draft passes.
 61- Wherever a mistake is expensive or irreversible: a person in the loop.
 62- Whatever you pick, invest in the check before the retry. With a check
 63  that catches every failure and one retry, a 95% step becomes 99.75%
 64  and ten steps succeed 97.5% of the time; with a check that catches half,
 65  77%; with no check, 60%. Retries are only as good as the check that
 66  triggers them.
 67
 68**What it costs.** Planning costs one extra model call up front and one per
 69replan. Checks cost engineering: a test suite, a schema, a query against
 70the source of truth, sometimes a second model with a rubric. Retries cost a
 71call each, but only for failures that were noticed. Checkpoints cost
 72storage and a small amount of bookkeeping, and they are what stop the
 73cost of a long task from growing exponentially: without them every extra
 74step multiplies the expected work by about 1.05, which is a straight line
 75on a log scale. Human review costs latency. Against all of this, the cost
 76of doing nothing is a task that is abandoned and restarted, paying for the
 77same steps over and over.
 78
 79**What breaks.**
 80
 81- **A plan reality has invalidated.** The API changed, the file moved, the
 82  assumption was wrong. Treat the plan as a hypothesis and the first
 83  failed step as evidence; replan from there, keeping what is done.
 84- **Self-review that approves the bug.** The reviewer shares the author's
 85  blind spots and can confidently reinforce a wrong answer. Put the check
 86  outside the model.
 87- **Retrying a failure nobody detected.** A silent failure gets no retry
 88  and poisons every later step. Measure your check's detection rate; it
 89  matters more than the retry count.
 90- **Restarting from step one.** Without checkpoints a late failure throws
 91  away all the work, and long runs become both unreliable and expensive.
 92  Save every output; rerun only the failed step.
 93- **A planner that loops.** Cap the replans and the total steps, and
 94  report the failure instead of trying forever.
 95- **Subtasks that cannot be checked.** "Reconcile the invoices" has no
 96  definition of done; "each invoice appears exactly once in the match" does.
 97  Make every subtask small, concrete and verifiable.
 98
 99**In the wild.** Wang et al. (2023), *Plan-and-Solve Prompting*, showed
100that asking a model to devise a plan and then carry it out cuts the
101missing-step errors of reasoning straight through. Shinn et al. (2023),
102*Reflexion*, turn external feedback such as failed tests into written
103lessons kept in memory for the next attempt, and report 91% pass@1 on
104HumanEval against 80% for the base model; the gains come from signals
105outside the model. Anthropic's *Building effective agents* names the
106workflow patterns in the table (prompt chaining, routing, parallelization,
107orchestrator-workers, evaluator-optimizer) and insists on stopping
108conditions such as a maximum number of iterations. Temporal is durable
109execution as a product: it persists a complete event history, and when a
110process dies another rebuilds the state and resumes where it stopped, with
111local variables intact. LangGraph's persistence and checkpoints are the
112same idea inside an agent framework.
113
114**Go deeper.** Level 2 builds the arithmetic and the remedies in plain
115Python: the compounding formula and its curves, a plan-and-execute loop
116that replans from a failed step, five subtasks with checks and the retry
117that reruns only one of them, the expected-cost formulas with and without
118checkpoints, and a reflection-versus-verification experiment on a
119leap-year function. If you only needed to choose, you are done.
120
121## Level 2: How it works, from scratch
122
123What follows measures the arithmetic, then builds the three remedies one
124mechanism at a time.
125
126An agent that does one thing is easy. An agent that does twenty things in a
127row mostly fails, for a reason that is pure arithmetic. This lesson measures
128that arithmetic, then builds the three remedies: **decompose** the task into
129steps you can check, **verify** each step with something outside the model,
130and **checkpoint** so a failure costs one step instead of the whole run.
131
132## 1. Compounding error: why long tasks fail
133
134**Everyday picture.** A relay race with ten runners. Each runner drops the
135baton only 1 time in 20 during their leg. That sounds safe, but the team
136needs *all ten* legs to go cleanly, and it loses about 4 races in 10.
137
138**Tiny worked example.** Each step of an agent succeeds 95% of the time.
139
140| steps | chance all succeed |
141|---|---|
142| 1 | 0.95 |
143| 2 | 0.95 × 0.95 = 0.9025 |
144| 10 | 0.95¹⁰ ≈ 0.599 |
145| 20 | 0.95²⁰ ≈ 0.358 |
146
147$$
148P(\text{task succeeds}) = p^{\,n}
149$$
150
151**Symbols**
152
153| Symbol | Meaning |
154|---|---|
155| $p$ | chance a single step succeeds |
156| $n$ | number of steps the task needs |
157
158**In words:** multiply the per-step success rate by itself once per step.
159
160**On the example:** $0.95^{10} \approx 0.60$ and $0.95^{20} \approx 0.36$.
161A 95%-reliable step is a 60%-reliable 10-step task.
162
163**In Python:**
164
165```python
166p = 0.95
167# ten steps that must all succeed
168round(p ** 10, 2)  # → 0.6
169round(p ** 20, 2)  # → 0.36
170```
171
172![At 95% per step a 10-step task succeeds only about 60% of the time, and every curve keeps sliding toward zero as steps are added](figures/primer.agents.planning.compounding.svg)
173
174**Reading it:** the x-axis is the number of steps in the task, and the y-axis is the
175chance the whole task succeeds. Every curve starts near the top and slides
176down, and the lower the per-step rate, the faster it falls. Find 10 steps on the
177x-axis and read up to the 95% curve: about 0.6. That's why a demo of a
178three-step task looks great and the same agent on a real twenty-step job
179disappoints.
180
181**In code:** `end_to_end_success` multiplies the per-step rate by itself once
182per step, which is the whole formula above.
183
184**Why it matters.** The remedies all attack $p$ or $n$: fewer steps
185(higher-level tools, see `primer.agents.tools`), checks that catch failures,
186retries of just the failed step, and human review at critical points.
187
188## 2. Plan-and-execute, with replanning
189
190**Everyday picture.** Cooking from a recipe. You read the whole recipe first
191(the *plan*), then do the steps. Halfway through you discover you're out of
192eggs. You don't start over, and you don't pretend you have eggs. You revise
193the rest of the recipe with what you've got, keeping everything you've
194already chopped.
195
196**Tiny worked example.** Goal "Reconcile Q3 invoices". The planner's first
197plan uses an old API. Here's the actual run:
198
199```text
200planner  -> {"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}
201execute  fetch_payments      ok
202execute  fetch_invoices_v1   failed: invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2
203planner  <- "Completed so far: fetch_payments. Step fetch_invoices_v1 failed: ...retired..."
204planner  -> {"steps": ["fetch_invoices_v2", "match", "draft_summary"]}
205execute  fetch_invoices_v2   ok
206execute  match               ok
207execute  draft_summary       ok     (fetch_payments was NOT run again)
208```
209
210```mermaid
211flowchart TD
212  G[Goal] --> P[Planner writes a plan]
213  P --> X[Execute next step]
214  X --> C{Step result<br/>meets expectations?}
215  C -->|yes| M{More steps?}
216  M -->|yes| X
217  M -->|no| D[Done]
218  C -->|no| R{Replans left?}
219  R -->|yes| F[Tell the planner:<br/>what's done + why it failed]
220  F --> P
221  R -->|no| S[Stop: report the failure]
222```
223
224**Reading it:** the inner loop (execute → check → next) is the plan being
225followed. The outer loop back to *Planner* only happens on a failed check,
226and it carries two facts: what's already done (so work is kept) and why the
227step failed (so the new plan can route around it). The *Replans left?*
228diamond stops a planner that can't adapt from looping forever.
229
230**The code.** `plan_and_execute(planner, goal)` asks for a JSON plan, runs
231each action, records outputs in `state`, and on `StepFailed` asks for a
232revised plan. Steps already in `state` are skipped.
233
234**In code:** `PlanRun` is what `plan_and_execute` returns: every plan the
235planner wrote, each executed step with "ok" or its failure, and the replan count.
236
237**Why it matters.** Deciding one step at a time (a pure ReAct loop, see
238`primer.agents.agent_loop`) drifts on long tasks. An upfront plan keeps the
239agent on track and lets a person see its intent before it acts. The risk is
240following a plan that reality has invalidated, which is why replanning exists.
241
242## 3. Decomposition into verifiable subtasks
243
244**Everyday picture.** Moving house. "Move house" isn't something you can check off,
245but "every box labelled", "van booked for Saturday" and "keys handed back"
246are. Each has a clear definition of done.
247
248**Tiny worked example.** "Reconcile Q3 invoices" becomes five subtasks, each
249with a check:
250
251| subtask | output | check (definition of done) |
252|---|---|---|
253| fetch_invoices | 4 invoices | at least one invoice returned |
254| fetch_payments | 3 payments | at least one payment returned |
255| match | 4 rows: invoiced vs paid | each invoice appears exactly once |
256| list_mismatches | INV-102 (1200 vs 1100), INV-104 (450 vs 0) | none |
257| draft_summary | "Q3 reconciliation: 4 invoices, 3 payments, 2 mismatches (INV-102 short by 100, INV-104 unpaid)." | mentions the mismatches |
258
259When the payments API times out once, **only `fetch_payments` runs again**:
260attempts are `{fetch_invoices: 1, fetch_payments: 2, match: 1, ...}`.
261
262```mermaid
263flowchart LR
264  A[fetch_invoices] --> C[match]
265  B[fetch_payments] --> C
266  C --> D[list_mismatches]
267  D --> E[draft_summary]
268  B -. timeout .-> B
269```
270
271**Reading it:** arrows show which outputs feed which step. The dotted
272self-loop on `fetch_payments` is a retry. Because every step's output is
273kept (a *checkpoint*), the retry doesn't redo `fetch_invoices`.
274
275**In code:** `Subtask` pairs a step's work with its check (the definition of
276done); `run_subtasks` runs them in order, retries only the step whose check
277failed, and returns a `RunReport` of outputs and attempts. `reconcile_q3`
278builds the five subtasks in the table.
279
280### Checkpoints: how much work a failure costs
281
282Without checkpoints, a failure anywhere means starting again from step one.
283With them, you retry just the failed step. The expected number of step
284executions to finish an $n$-step task:
285
286$$
287\text{with checkpoints: } \frac{n}{p}
288\qquad
289\text{restart from scratch: } \frac{1 - p^{n}}{(1-p)\,p^{n}}
290$$
291
292**Symbols**
293
294| Symbol | Meaning |
295|---|---|
296| $p$ | chance a single step attempt succeeds (failures are noticed) |
297| $n$ | number of steps |
298
299**In words:** with checkpoints each step needs on average $1/p$ attempts, so
300$n$ steps need $n/p$. Restarting from scratch needs $n$ successes *in a row*,
301and the expected wait for a run of $n$ successes grows roughly like $1/p^n$.
302
303**On the example:** $p = 0.95$, $n = 10$: with checkpoints
304$10 / 0.95 \approx 10.5$ step runs; restarting from scratch
305$(1 - 0.599) / (0.05 \times 0.599) \approx 13.4$. At $n = 50$ it's about
30652.6 vs. 240.
307
308**In Python:**
309
310```python
311def with_checkpoints(n, p):
312    # each step needs 1/p attempts on average
313    return n / p
314def restart_from_scratch(n, p):
315    return (1 - p ** n) / ((1 - p) * p ** n)
316round(with_checkpoints(10, 0.95), 1), round(restart_from_scratch(10, 0.95), 1)  # → (10.5, 13.4)
317round(with_checkpoints(50, 0.95), 1), round(restart_from_scratch(50, 0.95))  # → (52.6, 240)
318```
319
320![On the log scale, restarting from scratch climbs as a straight line (exponential growth) while checkpoints stay close to one run per step](figures/primer.agents.planning.checkpoints.svg)
321
322**Reading it:** the x-axis is task length, and the y-axis (log scale) is how
323many step executions you should expect to pay for. On a log scale each
324gridline is ten times the one below, so exponential growth draws a
325*straight* line. The restart line is that straight line: every extra step
326multiplies its cost by about $1/0.95 \approx 1.05$. The checkpoint line
327($n/p$, just proportional to $n$) bends over and flattens, because on a log
328scale going from 10 to 20 steps rises no more than going from 5 to 10. The
329widening gap between them is the cost of having no checkpoints, and at 50
330steps it's already about 240 step runs against 53. Long tasks without
331checkpoints aren't just unreliable; they're expensive. **Durable execution** is the engineering name
332for this: a workflow engine that saves each step's result so a crashed run
333resumes where it stopped.
334
335**In code:** `expected_step_runs` evaluates both formulas: $n/p$ with
336checkpoints, the restart formula without.
337
338## 4. Reflection vs. external verification
339
340**Everyday picture.** Proofreading your own essay versus having someone run the
341numbers in it. You read what you *meant* to write, so your own blind spots
342stay blind. A calculator doesn't share them.
343
344**Tiny worked example.** The model writes a leap-year function:
345
346```python
347def is_leap(year):
348    return year % 4 == 0        # forgets that 1900 was not a leap year
349```
350
351*Reflection* (asking a model to review it) replies "Looks correct: years
352divisible by 4 are leap years." *External verification* (running it against
353known answers) replies `is_leap(1900) returned True, expected False`. Fed
354that failure, the model's second draft passes every case.
355
356```mermaid
357flowchart LR
358  G[Model writes code] --> R1[Reflection:<br/>model reads its own code]
359  R1 -->|same blind spot| A1[Approved, still wrong]
360  G --> V[Verification:<br/>run tests / check schema / query the DB]
361  V -->|1900 fails| F[Failure fed back]
362  F --> G2[Second draft]
363  G2 --> V2[Checks pass]
364```
365
366**Reading it:** two paths leave the same first draft. The top path stays inside
367the model and ends at "approved, still wrong". The bottom path goes through
368something outside the model (tests, a schema validator, a database query),
369and that something produces a concrete, checkable failure the model can fix.
370
371**In code:** `self_review` is the top path (a model reads the code and
372approves or not); `run_checks` is the bottom path (it runs the code against
373known answers); `generate_until_checks_pass` loops the bottom path, feeding
374each failure back to the model until the checks pass.
375
376With verification and retries, the per-step success rate rises:
377
378$$
379p_{\text{step}} = p \sum_{i=0}^{r} \big((1-p)\,d\big)^{i}
380$$
381
382**Symbols**
383
384| Symbol | Meaning |
385|---|---|
386| $p$ | chance one attempt succeeds |
387| $d$ | chance the check *detects* a failed attempt |
388| $r$ | retries allowed after a detected failure |
389
390**In words:** you succeed on the first try, or fail *and notice* and succeed
391on the next try, and so on. Failures you don't notice get no retry.
392
393**On the example:** $p = 0.95$, perfect check $d = 1$, one retry:
394$0.95 + 0.05 \times 0.95 = 0.9975$ per step, so ten steps succeed
395$0.9975^{10} \approx 97.5\%$ of the time instead of 60%. With a check that
396catches only half the failures ($d = 0.5$): $0.97375^{10} \approx 77\%$.
397
398**In Python:**
399
400```python
401def p_step(p, d, r):
402    # p Σ_{i=0}^{r} ((1 - p) d)^i
403    return p * sum(((1 - p) * d) ** i for i in range(r + 1))
404# a perfect check, one retry
405round(p_step(0.95, 1, 1), 4)  # → 0.9975
406# ten such steps
407round(p_step(0.95, 1, 1) ** 10, 3)  # → 0.975
408# a check that catches half the failures
409round(p_step(0.95, 0.5, 1), 5)  # → 0.97375
410round(p_step(0.95, 0.5, 1) ** 10, 2)  # → 0.77
411```
412
413![With the same 95% step, a ten-step task succeeds 97.5% of the time when a check catches every failure and allows one retry, 77% when the check catches half, and 60% with no checks](figures/primer.agents.planning.verification.svg)
414
415**Reading it:** all curves use the same 95%-reliable step. The bottom curve has
416no checks. The middle ones add a retry after a failure is caught, with a
417check that catches half or all failures. The top curve allows two retries.
418The gap between "half" and "all" is the lesson: retries are only as good
419as the check that triggers them. Invest in the check.
420
421**In code:** `step_success` computes $p_{\text{step}}$ from $p$, $d$ and $r$;
422feed its result into `end_to_end_success` to get the ten-step curves.
423
424## In 20 seconds
425- Success compounds: $p^n$. 95% per step is 60% over ten steps.
426- Plan first, execute step by step, replan from the point of failure, and keep finished work.
427- Decompose into subtasks with explicit checks; retry only the failed step (checkpoints).
428- Self-review shares the model's blind spots; tests, schemas and database checks don't.
429- Retries help only for failures you detect, so the check matters more than the retry.
430
431## Self-test questions
432
433**Q: Each step of your agent is 95% reliable. How reliable is a 20-step task, and what do you do about it?**
434A: $0.95^{20} \approx 36\%$. Cut the steps (higher-level tools), add
435external checks after key steps with a retry of just that step, checkpoint
436so failures don't restart the run, and put human review where a mistake is
437expensive. With perfect checks and one retry, per-step reliability becomes
43899.75%, and 20 steps succeed about 95% of the time.
439
440**Q: Why isn't "ask the model to double-check its work" enough?**
441A: The reviewer shares the author's blind spots, and self-critique can
442confidently reinforce a wrong answer. Checks outside the model (tests,
443schema validation, a query against the source of truth, a separate grader
444model with a rubric) produce concrete failures the model can act on.
445
446**Q: When is plan-and-execute better than deciding one step at a time?**
447A: For long, multi-part tasks where staying on track matters and where
448showing the plan to a person before acting is valuable. Always pair it with
449replanning. A plan is a hypothesis, and the first failed step is evidence.
450
451**Q: What makes a subtask "good"?**
452A: It's small, concrete and verifiable. It has a definition of done that code
453can check, and its output is saved so a later failure doesn't redo it.
454
455## The papers behind this lesson
456
457- **Wang et al., *Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models* (2023).**
458  https://arxiv.org/abs/2305.04091. It showed that asking a model to first
459  devise a plan and then carry it out step by step reduces missed steps
460  compared with reasoning straight through.
461- **Shinn et al., *Reflexion: Language Agents with Verbal Reinforcement Learning* (2023).**
462  https://arxiv.org/abs/2303.11366. Agents that turn feedback from the
463  environment (failed tests, wrong answers) into written lessons and retry
464  improve markedly, and the gains come from *external* signals.
465  [annotated companion](../../papers/reflexion.html)
466
467## Further reading
468- Anthropic, *Building effective agents*: https://www.anthropic.com/engineering/building-effective-agents
469- Lilian Weng, *LLM Powered Autonomous Agents* (planning, reflection): https://lilianweng.github.io/posts/2023-06-23-agent/
470- Temporal, durable execution: https://docs.temporal.io/
471- LangGraph docs (persistence and checkpoints): https://langchain-ai.github.io/langgraph/
472"""
473
474from __future__ import annotations
475
476import json
477from dataclasses import dataclass, field
478from typing import Any, Callable
479
480from primer.agents.llm import LLM, ScriptedLLM, last_user_text
481
482
483def end_to_end_success(p: float, n_steps: int) -> float:
484    """Chance that n independent steps, each succeeding with probability p, all succeed."""
485    return p**n_steps
486
487
488def step_success(p: float, detect: float, retries: int) -> float:
489    """Chance one step ends up correct when failures are caught with probability `detect`
490    and each caught failure is retried, up to `retries` times.
491
492    Attempt 1 succeeds with p. It fails *and is noticed* with (1-p)·detect, which
493    buys another attempt. A failure nobody notices counts as a silent failure.
494
495    ```text
496    P = p · (1 + q + q² + ... + q^retries),  where q = (1-p)·detect
497    ```
498    """
499    q = (1 - p) * detect
500    return p * sum(q**i for i in range(retries + 1))
501
502
503def expected_step_runs(p: float, n_steps: int, checkpoints: bool) -> float:
504    """Expected number of step executions to get n steps done, retrying until success.
505
506    With checkpoints, each step retries on its own: n / p (a geometric wait per step).
507    Without them, any failure throws away all progress, so you need n successes
508    *in a row*: (1 - p^n) / ((1 - p) · p^n).
509    """
510    if checkpoints:
511        return n_steps / p
512    return (1 - p**n_steps) / ((1 - p) * p**n_steps)
513
514
515# ---------------------------------------------------------------------------
516# Decomposition: small, verifiable subtasks with checks and checkpoints
517# ---------------------------------------------------------------------------
518
519
520@dataclass
521class Subtask:
522    """One step of a decomposed task.
523
524    `run(outputs)` receives the outputs of earlier steps (by name) and returns
525    this step's output. `check(output)` returns a problem description, or None
526    when the output is acceptable. The check is the step's definition of done.
527    """
528
529    name: str
530    run: Callable[[dict[str, Any]], Any]
531    check: Callable[[Any], str | None] = lambda out: None
532
533
534@dataclass
535class RunReport:
536    status: str  # "done" | "failed"
537    outputs: dict[str, Any]
538    attempts: dict[str, int]
539    failed_step: str | None = None
540    problem: str | None = None
541
542
543def run_subtasks(subtasks: list[Subtask], max_retries: int = 2) -> RunReport:
544    """Run subtasks in order, checking each one and retrying only the one that failed."""
545    outputs: dict[str, Any] = {}
546    attempts: dict[str, int] = {}
547    for task in subtasks:
548        problem = None
549        for _ in range(1 + max_retries):
550            attempts[task.name] = attempts.get(task.name, 0) + 1
551            try:
552                out = task.run(outputs)
553                problem = task.check(out)
554            except Exception as e:  # noqa: BLE001  (a crash is just another failed attempt)
555                problem = f"{type(e).__name__}: {e}"
556            if problem is None:
557                outputs[task.name] = out
558                break
559        if problem is not None:
560            return RunReport("failed", outputs, attempts, task.name, problem)
561    return RunReport("done", outputs, attempts)
562
563
564INVOICES = [
565    {"invoice": "INV-101", "vendor": "Acme", "amount": 1500},
566    {"invoice": "INV-102", "vendor": "Globex", "amount": 1200},
567    {"invoice": "INV-103", "vendor": "Initech", "amount": 800},
568    {"invoice": "INV-104", "vendor": "Umbrella", "amount": 450},
569]
570PAYMENTS = [
571    {"payment": "PAY-9001", "invoice": "INV-101", "amount": 1500},
572    {"payment": "PAY-9002", "invoice": "INV-102", "amount": 1100},
573    {"payment": "PAY-9003", "invoice": "INV-103", "amount": 800},
574]
575
576
577def _match(outputs: dict[str, Any]) -> list[dict[str, Any]]:
578    paid: dict[str, int] = {}
579    for pay in outputs["fetch_payments"]:
580        paid[pay["invoice"]] = paid.get(pay["invoice"], 0) + pay["amount"]
581    return [
582        {"invoice": inv["invoice"], "vendor": inv["vendor"], "invoiced": inv["amount"], "paid": paid.get(inv["invoice"], 0)}
583        for inv in outputs["fetch_invoices"]
584    ]
585
586
587def _summary(outputs: dict[str, Any]) -> str:
588    notes = [
589        f"{m['invoice']} unpaid" if m["paid"] == 0 else f"{m['invoice']} short by {m['invoiced'] - m['paid']}"
590        for m in outputs["list_mismatches"]
591    ]
592    return (
593        f"Q3 reconciliation: {len(outputs['fetch_invoices'])} invoices, {len(outputs['fetch_payments'])} payments, "
594        f"{len(notes)} mismatches ({', '.join(notes)})."
595    )
596
597
598def reconcile_q3(
599    fetch_invoices: Callable[[], list[dict]] = lambda: list(INVOICES),
600    fetch_payments: Callable[[], list[dict]] = lambda: list(PAYMENTS),
601) -> list[Subtask]:
602    """"Reconcile Q3 invoices" decomposed into five steps, each with its own definition of done."""
603    return [
604        Subtask("fetch_invoices", lambda o: fetch_invoices(), lambda out: None if out else "no invoices returned for Q3"),
605        Subtask("fetch_payments", lambda o: fetch_payments(), lambda out: None if out else "no payments returned for Q3"),
606        # Each invoice must appear exactly once in the match, or rows were duplicated.
607        Subtask("match", _match, lambda out: None if len({m["invoice"] for m in out}) == len(out) else "duplicated invoices"),
608        Subtask("list_mismatches", lambda o: [m for m in o["match"] if m["paid"] != m["invoiced"]]),
609        Subtask("draft_summary", _summary, lambda out: None if "mismatch" in out else "summary does not report mismatches"),
610    ]
611
612
613# ---------------------------------------------------------------------------
614# Plan-and-execute with replanning
615# ---------------------------------------------------------------------------
616
617
618class StepFailed(Exception):
619    """A step's result didn't meet expectations. The message goes back to the planner."""
620
621
622def _fetch_invoices_v1(state: dict[str, Any]) -> list[dict]:
623    raise StepFailed("invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2")
624
625
626def _need(state: dict[str, Any], *keys: str) -> None:
627    missing = [k for k in keys if k not in state]
628    if missing:
629        raise StepFailed(f"cannot run yet: missing {missing}; fetch them first")
630
631
632def _match_step(state: dict[str, Any]) -> list[dict]:
633    _need(state, "fetch_invoices_v2", "fetch_payments")
634    return _match({"fetch_invoices": state["fetch_invoices_v2"], "fetch_payments": state["fetch_payments"]})
635
636
637def _summary_step(state: dict[str, Any]) -> str:
638    _need(state, "match")
639    mismatches = [m for m in state["match"] if m["paid"] != m["invoiced"]]
640    return _summary({"fetch_invoices": state["fetch_invoices_v2"], "fetch_payments": state["fetch_payments"], "list_mismatches": mismatches})
641
642
643# action name -> function(state) -> output (stored in state under the action name)
644PLAN_ACTIONS: dict[str, Callable[[dict[str, Any]], Any]] = {
645    "fetch_invoices_v1": _fetch_invoices_v1,
646    "fetch_invoices_v2": lambda state: list(INVOICES),
647    "fetch_payments": lambda state: list(PAYMENTS),
648    "match": _match_step,
649    "draft_summary": _summary_step,
650}
651
652
653@dataclass
654class PlanRun:
655    status: str  # "done" | "failed"
656    replans: int
657    executed: list[tuple[str, str]]  # (action, "ok" or "failed: why")
658    plans: list[list[str]]
659    state: dict[str, Any] = field(default_factory=dict)
660
661
662def _ask_planner(planner: LLM, goal: str, actions: list[str], done: list[str], feedback: str) -> list[str]:
663    prompt = (
664        f"Goal: {goal}\n"
665        f"Available actions: {', '.join(actions)}\n"
666        f"Completed so far: {', '.join(done) or 'nothing'}\n"
667        + (f"{feedback}\n" if feedback else "")
668        + 'Reply with JSON {"steps": [...]} listing the remaining actions in order.'
669    )
670    resp = planner.complete(system="You are a planner. Output only JSON.", messages=[{"role": "user", "content": prompt}])
671    return list(json.loads(resp.text)["steps"])
672
673
674def plan_and_execute(
675    planner: LLM,
676    goal: str,
677    actions: dict[str, Callable[[dict[str, Any]], Any]] | None = None,
678    max_replans: int = 2,
679) -> PlanRun:
680    """Ask for a plan, execute it step by step, and replan from where it broke.
681
682    Completed steps are remembered in `state`, so a revised plan never redoes
683    finished work. The planner only sees what it needs: the goal, the
684    actions, what's done, and why the last step failed.
685    """
686    actions = actions or PLAN_ACTIONS
687    state: dict[str, Any] = {}
688    executed: list[tuple[str, str]] = []
689    plans: list[list[str]] = []
690    replans, feedback = 0, ""
691    while True:
692        plan = _ask_planner(planner, goal, list(actions), [a for a in state], feedback)
693        plans.append(plan)
694        failure = None
695        for action in plan:
696            if action in state:
697                continue  # already done in an earlier plan: keep the work
698            try:
699                if action not in actions:
700                    raise StepFailed(f"unknown action {action!r}")
701                state[action] = actions[action](state)
702                executed.append((action, "ok"))
703            except StepFailed as e:
704                executed.append((action, f"failed: {e}"))
705                failure = f"Step {action} failed: {e}"
706                break
707        if failure is None:
708            return PlanRun("done", replans, executed, plans, state)
709        if replans >= max_replans:
710            return PlanRun("failed", replans, executed, plans, state)
711        replans += 1
712        feedback = failure
713
714
715def demo_planner() -> ScriptedLLM:
716    """A scripted planner that starts with the old API and adapts when told it's retired."""
717
718    def policy(system, messages, tools):
719        if "retired" in last_user_text(messages):
720            return '{"steps": ["fetch_invoices_v2", "match", "draft_summary"]}'
721        return '{"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}'
722
723    return ScriptedLLM(policy)
724
725
726# ---------------------------------------------------------------------------
727# Reflection vs. external verification
728# ---------------------------------------------------------------------------
729
730BUGGY_LEAP = "def is_leap(year):\n    return year % 4 == 0\n"
731FIXED_LEAP = "def is_leap(year):\n    return year % 4 == 0 and (year % 100 != 0 or year % 400 == 0)\n"
732
733# (year, is it a leap year?) The century cases are the ones a quick read misses.
734LEAP_CASES = [(2024, True), (2023, False), (1900, False), (2000, True)]
735
736
737def self_review(critic: LLM, code: str) -> tuple[bool, str]:
738    """Ask a model to review code by reading it. Returns (approved, its note)."""
739    resp = critic.complete(
740        system="Review the code. Start your reply with 'Looks correct' or 'Bug:'.",
741        messages=[{"role": "user", "content": code}],
742    )
743    return resp.text.startswith("Looks correct"), resp.text
744
745
746def run_checks(code: str) -> list[str]:
747    """External verification: actually run the code against known answers."""
748    namespace: dict[str, Any] = {}
749    exec(code, namespace)  # our own generated snippet, run in an isolated namespace
750    fn = namespace["is_leap"]
751    return [f"is_leap({y}) returned {fn(y)}, expected {want}" for y, want in LEAP_CASES if fn(y) != want]
752
753
754def confident_critic() -> ScriptedLLM:
755    """A reviewer that shares the author's blind spot, which is the usual failure of self-critique."""
756    return ScriptedLLM(["Looks correct: years divisible by 4 are leap years."])
757
758
759def fixing_coder() -> ScriptedLLM:
760    """Writes the buggy version first, then fixes it once shown a failing case."""
761
762    def policy(system, messages, tools):
763        return FIXED_LEAP if "1900" in last_user_text(messages) else BUGGY_LEAP
764
765    return ScriptedLLM(policy)
766
767
768def generate_until_checks_pass(coder: LLM, task: str, max_attempts: int = 3) -> tuple[str, int]:
769    """Generate, verify externally, feed failures back, repeat. Returns (code, attempts)."""
770    messages: list[dict[str, Any]] = [{"role": "user", "content": task}]
771    code = ""
772    for attempt in range(1, max_attempts + 1):
773        code = coder.complete(system="Reply with Python code only.", messages=messages).text
774        failures = run_checks(code)
775        if not failures:
776            return code, attempt
777        messages += [
778            {"role": "assistant", "content": code},
779            {"role": "user", "content": "These checks failed:\n" + "\n".join(failures)},
780        ]
781    return code, max_attempts
782
783
784# ---------------------------------------------------------------------------
785# Figures and walkthrough
786# ---------------------------------------------------------------------------
787
788
789def figures() -> dict:
790    """Plots computed from this lesson's own formulas (matplotlib imported here)."""
791    import matplotlib
792
793    matplotlib.use("Agg")
794    import matplotlib.pyplot as plt
795
796    figs = {}
797    steps = list(range(1, 31))
798
799    fig, ax = plt.subplots(figsize=(7, 4))
800    for p in (0.90, 0.95, 0.99):
801        ax.plot(steps, [end_to_end_success(p, n) for n in steps], "o-", ms=3, label=f"{p:.0%} per step")
802    ax.axvline(10, ls=":", color="gray")
803    ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds) = p^n", ylim=(0, 1.02),
804           title="Compounding error: reliable steps, unreliable tasks")
805    ax.legend()
806    figs["compounding"] = fig
807
808    fig, ax = plt.subplots(figsize=(7, 4))
809    variants = [
810        ("no checks", 0.95),
811        ("check catches half, 1 retry", step_success(0.95, 0.5, 1)),
812        ("check catches all, 1 retry", step_success(0.95, 1.0, 1)),
813        ("check catches all, 2 retries", step_success(0.95, 1.0, 2)),
814    ]
815    for label, ps in variants:
816        ax.plot(steps, [end_to_end_success(ps, n) for n in steps], label=f"{label} (step = {ps:.4f})")
817    ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds)", ylim=(0, 1.02),
818           title="Same 95% step, different verification")
819    ax.legend(fontsize=8)
820    figs["verification"] = fig
821
822    fig, ax = plt.subplots(figsize=(7, 4))
823    ns = list(range(1, 61))
824    ax.plot(ns, [expected_step_runs(0.95, n, True) for n in ns], label="checkpoints: retry the failed step (n/p)")
825    ax.plot(ns, [expected_step_runs(0.95, n, False) for n in ns], label="no checkpoints: restart from step 1")
826    ax.set_yscale("log")
827    ax.set(xlabel="steps in the task (n)", ylabel="expected step executions (log scale)",
828           title="What a failure costs (p = 0.95 per step)")
829    ax.legend()
830    figs["checkpoints"] = fig
831    return figs
832
833
834def demo() -> None:
835    from primer._show import banner, say, table, takeaway
836
837    banner("1. Compounding error")
838    table(["steps", "90%/step", "95%/step", "99%/step"],
839          [(n, *(end_to_end_success(p, n) for p in (0.90, 0.95, 0.99))) for n in (1, 5, 10, 20, 50)], floatfmt=".3f")
840    takeaway("A 95%-reliable step is a 60%-reliable ten-step task.")
841
842    banner("2. Plan-and-execute with replanning")
843    run = plan_and_execute(demo_planner(), "Reconcile Q3 invoices")
844    for i, plan in enumerate(run.plans):
845        print(f"  plan {i + 1}: {plan}")
846    print()
847    for action, result in run.executed:
848        print(f"  {action:18} {result}")
849    print()
850    say(f"status={run.status}, replans={run.replans}. fetch_payments ran once: finished work is kept.")
851    say(f"Summary: {run.state['draft_summary']}")
852
853    banner("3. Decomposition: retry only the failed step")
854    calls: list[int] = []
855
856    def flaky_payments() -> list[dict]:
857        calls.append(1)
858        if len(calls) == 1:
859            raise TimeoutError("payments API timed out")
860        return list(PAYMENTS)
861
862    report = run_subtasks(reconcile_q3(fetch_payments=flaky_payments))
863    table(["subtask", "attempts"], list(report.attempts.items()))
864    for m in report.outputs["list_mismatches"]:
865        print(f"  mismatch: {m}")
866    print()
867    table(["steps", "step runs with checkpoints", "step runs restarting"],
868          [(n, expected_step_runs(0.95, n, True), expected_step_runs(0.95, n, False)) for n in (5, 10, 20, 50)], floatfmt=".1f")
869
870    banner("4. Reflection vs. external verification")
871    approved, note = self_review(confident_critic(), BUGGY_LEAP)
872    print(BUGGY_LEAP)
873    say(f"Self-review: approved={approved}: {note!r}")
874    say(f"Running the test cases: {run_checks(BUGGY_LEAP)}")
875    code, attempts = generate_until_checks_pass(fixing_coder(), "Write is_leap(year).")
876    say(f"Feeding the failure back: passes on attempt {attempts}.")
877    table(["verification", "per-step", "10-step task"],
878          [(label, ps, end_to_end_success(ps, 10)) for label, ps in (
879              ("none", 0.95), ("catches half, 1 retry", step_success(0.95, 0.5, 1)),
880              ("catches all, 1 retry", step_success(0.95, 1.0, 1)))], floatfmt=".4f")
881    takeaway("Retries are only as good as the check that triggers them. Check outside the model.")
882
883
884if __name__ == "__main__":
885    demo()
Level 3: the code, function by function.
def end_to_end_success(p: float, n_steps: int) -> float: on GitHub
484def end_to_end_success(p: float, n_steps: int) -> float:
485    """Chance that n independent steps, each succeeding with probability p, all succeed."""
486    return p**n_steps

Chance that n independent steps, each succeeding with probability p, all succeed.

def step_success(p: float, detect: float, retries: int) -> float: on GitHub
489def step_success(p: float, detect: float, retries: int) -> float:
490    """Chance one step ends up correct when failures are caught with probability `detect`
491    and each caught failure is retried, up to `retries` times.
492
493    Attempt 1 succeeds with p. It fails *and is noticed* with (1-p)·detect, which
494    buys another attempt. A failure nobody notices counts as a silent failure.
495
496    ```text
497    P = p · (1 + q + q² + ... + q^retries),  where q = (1-p)·detect
498    ```
499    """
500    q = (1 - p) * detect
501    return p * sum(q**i for i in range(retries + 1))

Chance one step ends up correct when failures are caught with probability detect and each caught failure is retried, up to retries times.

Attempt 1 succeeds with p. It fails and is noticed with (1-p)·detect, which buys another attempt. A failure nobody notices counts as a silent failure.

P = p · (1 + q + q² + ... + q^retries),  where q = (1-p)·detect
def expected_step_runs(p: float, n_steps: int, checkpoints: bool) -> float: on GitHub
504def expected_step_runs(p: float, n_steps: int, checkpoints: bool) -> float:
505    """Expected number of step executions to get n steps done, retrying until success.
506
507    With checkpoints, each step retries on its own: n / p (a geometric wait per step).
508    Without them, any failure throws away all progress, so you need n successes
509    *in a row*: (1 - p^n) / ((1 - p) · p^n).
510    """
511    if checkpoints:
512        return n_steps / p
513    return (1 - p**n_steps) / ((1 - p) * p**n_steps)

Expected number of step executions to get n steps done, retrying until success.

With checkpoints, each step retries on its own: n / p (a geometric wait per step). Without them, any failure throws away all progress, so you need n successes in a row: (1 - p^n) / ((1 - p) · p^n).

@dataclass
class Subtask: on GitHub
521@dataclass
522class Subtask:
523    """One step of a decomposed task.
524
525    `run(outputs)` receives the outputs of earlier steps (by name) and returns
526    this step's output. `check(output)` returns a problem description, or None
527    when the output is acceptable. The check is the step's definition of done.
528    """
529
530    name: str
531    run: Callable[[dict[str, Any]], Any]
532    check: Callable[[Any], str | None] = lambda out: None

One step of a decomposed task.

run(outputs) receives the outputs of earlier steps (by name) and returns this step's output. check(output) returns a problem description, or None when the output is acceptable. The check is the step's definition of done.

Subtask( name: str, run: Callable[[dict[str, Any]], Any], check: Callable[[Any], str | None] = <function Subtask.<lambda>>)
name: str
run: Callable[[dict[str, Any]], Any]
def check(out): on GitHub
532    check: Callable[[Any], str | None] = lambda out: None
@dataclass
class RunReport: on GitHub
535@dataclass
536class RunReport:
537    status: str  # "done" | "failed"
538    outputs: dict[str, Any]
539    attempts: dict[str, int]
540    failed_step: str | None = None
541    problem: str | None = None
RunReport( status: str, outputs: dict[str, typing.Any], attempts: dict[str, int], failed_step: str | None = None, problem: str | None = None)
status: str
outputs: dict[str, typing.Any]
attempts: dict[str, int]
failed_step: str | None = None
problem: str | None = None
def run_subtasks( subtasks: list[Subtask], max_retries: int = 2) -> RunReport: on GitHub
544def run_subtasks(subtasks: list[Subtask], max_retries: int = 2) -> RunReport:
545    """Run subtasks in order, checking each one and retrying only the one that failed."""
546    outputs: dict[str, Any] = {}
547    attempts: dict[str, int] = {}
548    for task in subtasks:
549        problem = None
550        for _ in range(1 + max_retries):
551            attempts[task.name] = attempts.get(task.name, 0) + 1
552            try:
553                out = task.run(outputs)
554                problem = task.check(out)
555            except Exception as e:  # noqa: BLE001  (a crash is just another failed attempt)
556                problem = f"{type(e).__name__}: {e}"
557            if problem is None:
558                outputs[task.name] = out
559                break
560        if problem is not None:
561            return RunReport("failed", outputs, attempts, task.name, problem)
562    return RunReport("done", outputs, attempts)

Run subtasks in order, checking each one and retrying only the one that failed.

INVOICES = [{'invoice': 'INV-101', 'vendor': 'Acme', 'amount': 1500}, {'invoice': 'INV-102', 'vendor': 'Globex', 'amount': 1200}, {'invoice': 'INV-103', 'vendor': 'Initech', 'amount': 800}, {'invoice': 'INV-104', 'vendor': 'Umbrella', 'amount': 450}]
PAYMENTS = [{'payment': 'PAY-9001', 'invoice': 'INV-101', 'amount': 1500}, {'payment': 'PAY-9002', 'invoice': 'INV-102', 'amount': 1100}, {'payment': 'PAY-9003', 'invoice': 'INV-103', 'amount': 800}]
def reconcile_q3( fetch_invoices: Callable[[], list[dict]] = <function <lambda>>, fetch_payments: Callable[[], list[dict]] = <function <lambda>>) -> list[Subtask]: on GitHub
599def reconcile_q3(
600    fetch_invoices: Callable[[], list[dict]] = lambda: list(INVOICES),
601    fetch_payments: Callable[[], list[dict]] = lambda: list(PAYMENTS),
602) -> list[Subtask]:
603    """"Reconcile Q3 invoices" decomposed into five steps, each with its own definition of done."""
604    return [
605        Subtask("fetch_invoices", lambda o: fetch_invoices(), lambda out: None if out else "no invoices returned for Q3"),
606        Subtask("fetch_payments", lambda o: fetch_payments(), lambda out: None if out else "no payments returned for Q3"),
607        # Each invoice must appear exactly once in the match, or rows were duplicated.
608        Subtask("match", _match, lambda out: None if len({m["invoice"] for m in out}) == len(out) else "duplicated invoices"),
609        Subtask("list_mismatches", lambda o: [m for m in o["match"] if m["paid"] != m["invoiced"]]),
610        Subtask("draft_summary", _summary, lambda out: None if "mismatch" in out else "summary does not report mismatches"),
611    ]

"Reconcile Q3 invoices" decomposed into five steps, each with its own definition of done.

class StepFailed(builtins.Exception): on GitHub
619class StepFailed(Exception):
620    """A step's result didn't meet expectations. The message goes back to the planner."""

A step's result didn't meet expectations. The message goes back to the planner.

PLAN_ACTIONS: dict[str, typing.Callable[[dict[str, typing.Any]], typing.Any]] = {'fetch_invoices_v1': <function _fetch_invoices_v1>, 'fetch_invoices_v2': <function <lambda>>, 'fetch_payments': <function <lambda>>, 'match': <function _match_step>, 'draft_summary': <function _summary_step>}
@dataclass
class PlanRun: on GitHub
654@dataclass
655class PlanRun:
656    status: str  # "done" | "failed"
657    replans: int
658    executed: list[tuple[str, str]]  # (action, "ok" or "failed: why")
659    plans: list[list[str]]
660    state: dict[str, Any] = field(default_factory=dict)
PlanRun( status: str, replans: int, executed: list[tuple[str, str]], plans: list[list[str]], state: dict[str, typing.Any] = <factory>)
status: str
replans: int
executed: list[tuple[str, str]]
plans: list[list[str]]
state: dict[str, typing.Any]
def plan_and_execute( planner: primer.agents.llm.LLM, goal: str, actions: dict[str, typing.Callable[[dict[str, typing.Any]], typing.Any]] | None = None, max_replans: int = 2) -> PlanRun: on GitHub
675def plan_and_execute(
676    planner: LLM,
677    goal: str,
678    actions: dict[str, Callable[[dict[str, Any]], Any]] | None = None,
679    max_replans: int = 2,
680) -> PlanRun:
681    """Ask for a plan, execute it step by step, and replan from where it broke.
682
683    Completed steps are remembered in `state`, so a revised plan never redoes
684    finished work. The planner only sees what it needs: the goal, the
685    actions, what's done, and why the last step failed.
686    """
687    actions = actions or PLAN_ACTIONS
688    state: dict[str, Any] = {}
689    executed: list[tuple[str, str]] = []
690    plans: list[list[str]] = []
691    replans, feedback = 0, ""
692    while True:
693        plan = _ask_planner(planner, goal, list(actions), [a for a in state], feedback)
694        plans.append(plan)
695        failure = None
696        for action in plan:
697            if action in state:
698                continue  # already done in an earlier plan: keep the work
699            try:
700                if action not in actions:
701                    raise StepFailed(f"unknown action {action!r}")
702                state[action] = actions[action](state)
703                executed.append((action, "ok"))
704            except StepFailed as e:
705                executed.append((action, f"failed: {e}"))
706                failure = f"Step {action} failed: {e}"
707                break
708        if failure is None:
709            return PlanRun("done", replans, executed, plans, state)
710        if replans >= max_replans:
711            return PlanRun("failed", replans, executed, plans, state)
712        replans += 1
713        feedback = failure

Ask for a plan, execute it step by step, and replan from where it broke.

Completed steps are remembered in state, so a revised plan never redoes finished work. The planner only sees what it needs: the goal, the actions, what's done, and why the last step failed.

716def demo_planner() -> ScriptedLLM:
717    """A scripted planner that starts with the old API and adapts when told it's retired."""
718
719    def policy(system, messages, tools):
720        if "retired" in last_user_text(messages):
721            return '{"steps": ["fetch_invoices_v2", "match", "draft_summary"]}'
722        return '{"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}'
723
724    return ScriptedLLM(policy)

A scripted planner that starts with the old API and adapts when told it's retired.

BUGGY_LEAP = 'def is_leap(year):\n return year % 4 == 0\n'
FIXED_LEAP = 'def is_leap(year):\n return year % 4 == 0 and (year % 100 != 0 or year % 400 == 0)\n'
LEAP_CASES = [(2024, True), (2023, False), (1900, False), (2000, True)]
def self_review(critic: primer.agents.llm.LLM, code: str) -> tuple[bool, str]: on GitHub
738def self_review(critic: LLM, code: str) -> tuple[bool, str]:
739    """Ask a model to review code by reading it. Returns (approved, its note)."""
740    resp = critic.complete(
741        system="Review the code. Start your reply with 'Looks correct' or 'Bug:'.",
742        messages=[{"role": "user", "content": code}],
743    )
744    return resp.text.startswith("Looks correct"), resp.text

Ask a model to review code by reading it. Returns (approved, its note).

def run_checks(code: str) -> list[str]: on GitHub
747def run_checks(code: str) -> list[str]:
748    """External verification: actually run the code against known answers."""
749    namespace: dict[str, Any] = {}
750    exec(code, namespace)  # our own generated snippet, run in an isolated namespace
751    fn = namespace["is_leap"]
752    return [f"is_leap({y}) returned {fn(y)}, expected {want}" for y, want in LEAP_CASES if fn(y) != want]

External verification: actually run the code against known answers.

755def confident_critic() -> ScriptedLLM:
756    """A reviewer that shares the author's blind spot, which is the usual failure of self-critique."""
757    return ScriptedLLM(["Looks correct: years divisible by 4 are leap years."])

A reviewer that shares the author's blind spot, which is the usual failure of self-critique.

760def fixing_coder() -> ScriptedLLM:
761    """Writes the buggy version first, then fixes it once shown a failing case."""
762
763    def policy(system, messages, tools):
764        return FIXED_LEAP if "1900" in last_user_text(messages) else BUGGY_LEAP
765
766    return ScriptedLLM(policy)

Writes the buggy version first, then fixes it once shown a failing case.

def generate_until_checks_pass( coder: primer.agents.llm.LLM, task: str, max_attempts: int = 3) -> tuple[str, int]: on GitHub
769def generate_until_checks_pass(coder: LLM, task: str, max_attempts: int = 3) -> tuple[str, int]:
770    """Generate, verify externally, feed failures back, repeat. Returns (code, attempts)."""
771    messages: list[dict[str, Any]] = [{"role": "user", "content": task}]
772    code = ""
773    for attempt in range(1, max_attempts + 1):
774        code = coder.complete(system="Reply with Python code only.", messages=messages).text
775        failures = run_checks(code)
776        if not failures:
777            return code, attempt
778        messages += [
779            {"role": "assistant", "content": code},
780            {"role": "user", "content": "These checks failed:\n" + "\n".join(failures)},
781        ]
782    return code, max_attempts

Generate, verify externally, feed failures back, repeat. Returns (code, attempts).

def figures() -> dict: on GitHub
790def figures() -> dict:
791    """Plots computed from this lesson's own formulas (matplotlib imported here)."""
792    import matplotlib
793
794    matplotlib.use("Agg")
795    import matplotlib.pyplot as plt
796
797    figs = {}
798    steps = list(range(1, 31))
799
800    fig, ax = plt.subplots(figsize=(7, 4))
801    for p in (0.90, 0.95, 0.99):
802        ax.plot(steps, [end_to_end_success(p, n) for n in steps], "o-", ms=3, label=f"{p:.0%} per step")
803    ax.axvline(10, ls=":", color="gray")
804    ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds) = p^n", ylim=(0, 1.02),
805           title="Compounding error: reliable steps, unreliable tasks")
806    ax.legend()
807    figs["compounding"] = fig
808
809    fig, ax = plt.subplots(figsize=(7, 4))
810    variants = [
811        ("no checks", 0.95),
812        ("check catches half, 1 retry", step_success(0.95, 0.5, 1)),
813        ("check catches all, 1 retry", step_success(0.95, 1.0, 1)),
814        ("check catches all, 2 retries", step_success(0.95, 1.0, 2)),
815    ]
816    for label, ps in variants:
817        ax.plot(steps, [end_to_end_success(ps, n) for n in steps], label=f"{label} (step = {ps:.4f})")
818    ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds)", ylim=(0, 1.02),
819           title="Same 95% step, different verification")
820    ax.legend(fontsize=8)
821    figs["verification"] = fig
822
823    fig, ax = plt.subplots(figsize=(7, 4))
824    ns = list(range(1, 61))
825    ax.plot(ns, [expected_step_runs(0.95, n, True) for n in ns], label="checkpoints: retry the failed step (n/p)")
826    ax.plot(ns, [expected_step_runs(0.95, n, False) for n in ns], label="no checkpoints: restart from step 1")
827    ax.set_yscale("log")
828    ax.set(xlabel="steps in the task (n)", ylabel="expected step executions (log scale)",
829           title="What a failure costs (p = 0.95 per step)")
830    ax.legend()
831    figs["checkpoints"] = fig
832    return figs

Plots computed from this lesson's own formulas (matplotlib imported here).

def demo() -> None: on GitHub
835def demo() -> None:
836    from primer._show import banner, say, table, takeaway
837
838    banner("1. Compounding error")
839    table(["steps", "90%/step", "95%/step", "99%/step"],
840          [(n, *(end_to_end_success(p, n) for p in (0.90, 0.95, 0.99))) for n in (1, 5, 10, 20, 50)], floatfmt=".3f")
841    takeaway("A 95%-reliable step is a 60%-reliable ten-step task.")
842
843    banner("2. Plan-and-execute with replanning")
844    run = plan_and_execute(demo_planner(), "Reconcile Q3 invoices")
845    for i, plan in enumerate(run.plans):
846        print(f"  plan {i + 1}: {plan}")
847    print()
848    for action, result in run.executed:
849        print(f"  {action:18} {result}")
850    print()
851    say(f"status={run.status}, replans={run.replans}. fetch_payments ran once: finished work is kept.")
852    say(f"Summary: {run.state['draft_summary']}")
853
854    banner("3. Decomposition: retry only the failed step")
855    calls: list[int] = []
856
857    def flaky_payments() -> list[dict]:
858        calls.append(1)
859        if len(calls) == 1:
860            raise TimeoutError("payments API timed out")
861        return list(PAYMENTS)
862
863    report = run_subtasks(reconcile_q3(fetch_payments=flaky_payments))
864    table(["subtask", "attempts"], list(report.attempts.items()))
865    for m in report.outputs["list_mismatches"]:
866        print(f"  mismatch: {m}")
867    print()
868    table(["steps", "step runs with checkpoints", "step runs restarting"],
869          [(n, expected_step_runs(0.95, n, True), expected_step_runs(0.95, n, False)) for n in (5, 10, 20, 50)], floatfmt=".1f")
870
871    banner("4. Reflection vs. external verification")
872    approved, note = self_review(confident_critic(), BUGGY_LEAP)
873    print(BUGGY_LEAP)
874    say(f"Self-review: approved={approved}: {note!r}")
875    say(f"Running the test cases: {run_checks(BUGGY_LEAP)}")
876    code, attempts = generate_until_checks_pass(fixing_coder(), "Write is_leap(year).")
877    say(f"Feeding the failure back: passes on attempt {attempts}.")
878    table(["verification", "per-step", "10-step task"],
879          [(label, ps, end_to_end_success(ps, 10)) for label, ps in (
880              ("none", 0.95), ("catches half, 1 retry", step_success(0.95, 0.5, 1)),
881              ("catches all, 1 retry", step_success(0.95, 1.0, 1)))], floatfmt=".4f")
882    takeaway("Retries are only as good as the check that triggers them. Check outside the model.")