primer.agents.planning
Multi-step planning: plans, checks, retries and why long tasks fail
Run: python -m primer.agents.planning
This lesson builds on the step-at-a-time loop from
primer.agents.agent_loop, on tools from primer.agents.tools, and on
task state kept outside the model from primer.agents.memory.
Level 1: The practitioner's guide
In one sentence. Multi-step planning is how an agent gets a twenty-step job done when each step is only mostly reliable: write the steps down, check each one with something outside the model, save each result so a failure costs one step, and revise the plan from the point of failure rather than starting over.
When you need it. The moment a task takes more than a handful of model calls in a row. The arithmetic is unforgiving: a step that succeeds 95% of the time gives a ten-step task a 60% chance of finishing and a twenty-step task 36%; at fifty steps it is under 8%. Even a 99% step finishes a fifty-step task only 60% of the time. That is why a demo of a three-step task looks great and the same agent on a real twenty-step job disappoints. You don't need any of this for a task that is one or two calls long, or for a fixed sequence your code can run without asking a model what to do next. The tell: an agent that finishes the small cases and fails the long ones, and a bill that shows it restarting from step one after every stumble.
Your options. From the simplest to the most robust; each later row usually keeps the earlier ones:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| One step at a time | The model decides the next action after each result | Adapts to anything it sees | Drifts on long tasks; nobody sees the intent before it acts | The agent loop |
| A fixed workflow in code | Your code chains the calls in a known order (chaining, routing, parallel branches) | Predictable path, easy to test | Only for tasks whose steps you know in advance | Your code |
| Plan-and-execute with replanning | The model writes the whole plan first, executes it step by step, and rewrites the rest on a failure while keeping finished work | A plan a person can read before it runs; recovery that routes around the failure | An extra planning call per replan, and a cap on replans | Your loop around the model |
| Decomposition with checks | Each subtask has a definition of done that code can check | A failure is caught where it happens | Writing a check per subtask | Your code |
| Checkpoints | Every subtask's output is saved; a retry reruns only the failed step | A failure costs one step, not the run | Storage for each output | Your code, or a workflow engine |
| External verification with retries | Tests, a schema validator, a database query, or a separate grader with a rubric, fed back to the model | Concrete failures the model can fix; retries that actually help | The check itself, and one more call per retry | Outside the model |
| Human review at critical points | A person approves before an expensive or irreversible step | Mistakes stop before they cost | Latency and attention | Your product |
How to choose. Start with the simplest thing and add a row only when the numbers say so.
- Steps you can list in advance: a workflow in code. Anthropic's guidance is to find the simplest solution and add complexity only when needed, reserving agents for open-ended problems where the steps cannot be predicted.
- Steps the model must discover: plan-and-execute, with replanning and a replan limit so a planner that cannot adapt doesn't loop forever. In this lesson the first plan uses a retired API, the second routes around it, and the step that already succeeded is not run again.
- Any task longer than a few steps: decompose into subtasks with checks, and checkpoint. The expected work to finish a ten-step task at 95% per step is 10.5 step runs with checkpoints against 13.4 restarting from scratch; at fifty steps it is 53 against about 240.
- Wherever a wrong step can hide: verify outside the model. Asked to review its own leap-year function, this lesson's model approves the bug; a test against 1900 catches it, and the second draft passes.
- Wherever a mistake is expensive or irreversible: a person in the loop.
- Whatever you pick, invest in the check before the retry. With a check that catches every failure and one retry, a 95% step becomes 99.75% and ten steps succeed 97.5% of the time; with a check that catches half, 77%; with no check, 60%. Retries are only as good as the check that triggers them.
What it costs. Planning costs one extra model call up front and one per replan. Checks cost engineering: a test suite, a schema, a query against the source of truth, sometimes a second model with a rubric. Retries cost a call each, but only for failures that were noticed. Checkpoints cost storage and a small amount of bookkeeping, and they are what stop the cost of a long task from growing exponentially: without them every extra step multiplies the expected work by about 1.05, which is a straight line on a log scale. Human review costs latency. Against all of this, the cost of doing nothing is a task that is abandoned and restarted, paying for the same steps over and over.
What breaks.
- A plan reality has invalidated. The API changed, the file moved, the assumption was wrong. Treat the plan as a hypothesis and the first failed step as evidence; replan from there, keeping what is done.
- Self-review that approves the bug. The reviewer shares the author's blind spots and can confidently reinforce a wrong answer. Put the check outside the model.
- Retrying a failure nobody detected. A silent failure gets no retry and poisons every later step. Measure your check's detection rate; it matters more than the retry count.
- Restarting from step one. Without checkpoints a late failure throws away all the work, and long runs become both unreliable and expensive. Save every output; rerun only the failed step.
- A planner that loops. Cap the replans and the total steps, and report the failure instead of trying forever.
- Subtasks that cannot be checked. "Reconcile the invoices" has no definition of done; "each invoice appears exactly once in the match" does. Make every subtask small, concrete and verifiable.
In the wild. Wang et al. (2023), Plan-and-Solve Prompting, showed that asking a model to devise a plan and then carry it out cuts the missing-step errors of reasoning straight through. Shinn et al. (2023), Reflexion, turn external feedback such as failed tests into written lessons kept in memory for the next attempt, and report 91% pass@1 on HumanEval against 80% for the base model; the gains come from signals outside the model. Anthropic's Building effective agents names the workflow patterns in the table (prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer) and insists on stopping conditions such as a maximum number of iterations. Temporal is durable execution as a product: it persists a complete event history, and when a process dies another rebuilds the state and resumes where it stopped, with local variables intact. LangGraph's persistence and checkpoints are the same idea inside an agent framework.
Go deeper. Level 2 builds the arithmetic and the remedies in plain Python: the compounding formula and its curves, a plan-and-execute loop that replans from a failed step, five subtasks with checks and the retry that reruns only one of them, the expected-cost formulas with and without checkpoints, and a reflection-versus-verification experiment on a leap-year function. If you only needed to choose, you are done.
Level 2: How it works, from scratch
What follows measures the arithmetic, then builds the three remedies one mechanism at a time.
An agent that does one thing is easy. An agent that does twenty things in a row mostly fails, for a reason that is pure arithmetic. This lesson measures that arithmetic, then builds the three remedies: decompose the task into steps you can check, verify each step with something outside the model, and checkpoint so a failure costs one step instead of the whole run.
1. Compounding error: why long tasks fail
Everyday picture. A relay race with ten runners. Each runner drops the baton only 1 time in 20 during their leg. That sounds safe, but the team needs all ten legs to go cleanly, and it loses about 4 races in 10.
Tiny worked example. Each step of an agent succeeds 95% of the time.
| steps | chance all succeed |
|---|---|
| 1 | 0.95 |
| 2 | 0.95 × 0.95 = 0.9025 |
| 10 | 0.95¹⁰ ≈ 0.599 |
| 20 | 0.95²⁰ ≈ 0.358 |
Level 3: the formula and its symbols
$$ P(\text{task succeeds}) = p^{\,n} $$
Symbols
| Symbol | Meaning |
|---|---|
| $p$ | chance a single step succeeds |
| $n$ | number of steps the task needs |
In words: multiply the per-step success rate by itself once per step.
On the example: $0.95^{10} \approx 0.60$ and $0.95^{20} \approx 0.36$. A 95%-reliable step is a 60%-reliable 10-step task.
Level 3: in Python
In Python:
p = 0.95
# ten steps that must all succeed
round(p ** 10, 2) # → 0.6
round(p ** 20, 2) # → 0.36
Reading it: the x-axis is the number of steps in the task, and the y-axis is the chance the whole task succeeds. Every curve starts near the top and slides down, and the lower the per-step rate, the faster it falls. Find 10 steps on the x-axis and read up to the 95% curve: about 0.6. That's why a demo of a three-step task looks great and the same agent on a real twenty-step job disappoints.
In code: end_to_end_success multiplies the per-step rate by itself once
per step, which is the whole formula above.
Why it matters. The remedies all attack $p$ or $n$: fewer steps
(higher-level tools, see primer.agents.tools), checks that catch failures,
retries of just the failed step, and human review at critical points.
2. Plan-and-execute, with replanning
Everyday picture. Cooking from a recipe. You read the whole recipe first (the plan), then do the steps. Halfway through you discover you're out of eggs. You don't start over, and you don't pretend you have eggs. You revise the rest of the recipe with what you've got, keeping everything you've already chopped.
Tiny worked example. Goal "Reconcile Q3 invoices". The planner's first plan uses an old API. Here's the actual run:
planner -> {"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}
execute fetch_payments ok
execute fetch_invoices_v1 failed: invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2
planner <- "Completed so far: fetch_payments. Step fetch_invoices_v1 failed: ...retired..."
planner -> {"steps": ["fetch_invoices_v2", "match", "draft_summary"]}
execute fetch_invoices_v2 ok
execute match ok
execute draft_summary ok (fetch_payments was NOT run again)
flowchart TD G[Goal] --> P[Planner writes a plan] P --> X[Execute next step] X --> C{Step result<br/>meets expectations?} C -->|yes| M{More steps?} M -->|yes| X M -->|no| D[Done] C -->|no| R{Replans left?} R -->|yes| F[Tell the planner:<br/>what's done + why it failed] F --> P R -->|no| S[Stop: report the failure]
Reading it: the inner loop (execute → check → next) is the plan being followed. The outer loop back to Planner only happens on a failed check, and it carries two facts: what's already done (so work is kept) and why the step failed (so the new plan can route around it). The Replans left? diamond stops a planner that can't adapt from looping forever.
The code. plan_and_execute(planner, goal) asks for a JSON plan, runs
each action, records outputs in state, and on StepFailed asks for a
revised plan. Steps already in state are skipped.
In code: PlanRun is what plan_and_execute returns: every plan the
planner wrote, each executed step with "ok" or its failure, and the replan count.
Why it matters. Deciding one step at a time (a pure ReAct loop, see
primer.agents.agent_loop) drifts on long tasks. An upfront plan keeps the
agent on track and lets a person see its intent before it acts. The risk is
following a plan that reality has invalidated, which is why replanning exists.
3. Decomposition into verifiable subtasks
Everyday picture. Moving house. "Move house" isn't something you can check off, but "every box labelled", "van booked for Saturday" and "keys handed back" are. Each has a clear definition of done.
Tiny worked example. "Reconcile Q3 invoices" becomes five subtasks, each with a check:
| subtask | output | check (definition of done) |
|---|---|---|
| fetch_invoices | 4 invoices | at least one invoice returned |
| fetch_payments | 3 payments | at least one payment returned |
| match | 4 rows: invoiced vs paid | each invoice appears exactly once |
| list_mismatches | INV-102 (1200 vs 1100), INV-104 (450 vs 0) | none |
| draft_summary | "Q3 reconciliation: 4 invoices, 3 payments, 2 mismatches (INV-102 short by 100, INV-104 unpaid)." | mentions the mismatches |
When the payments API times out once, only fetch_payments runs again:
attempts are {fetch_invoices: 1, fetch_payments: 2, match: 1, ...}.
flowchart LR A[fetch_invoices] --> C[match] B[fetch_payments] --> C C --> D[list_mismatches] D --> E[draft_summary] B -. timeout .-> B
Reading it: arrows show which outputs feed which step. The dotted
self-loop on fetch_payments is a retry. Because every step's output is
kept (a checkpoint), the retry doesn't redo fetch_invoices.
In code: Subtask pairs a step's work with its check (the definition of
done); run_subtasks runs them in order, retries only the step whose check
failed, and returns a RunReport of outputs and attempts. reconcile_q3
builds the five subtasks in the table.
Checkpoints: how much work a failure costs
Without checkpoints, a failure anywhere means starting again from step one. With them, you retry just the failed step. The expected number of step executions to finish an $n$-step task:
Level 3: the formula and its symbols
$$ \text{with checkpoints: } \frac{n}{p} \qquad \text{restart from scratch: } \frac{1 - p^{n}}{(1-p)\,p^{n}} $$
Symbols
| Symbol | Meaning |
|---|---|
| $p$ | chance a single step attempt succeeds (failures are noticed) |
| $n$ | number of steps |
In words: with checkpoints each step needs on average $1/p$ attempts, so $n$ steps need $n/p$. Restarting from scratch needs $n$ successes in a row, and the expected wait for a run of $n$ successes grows roughly like $1/p^n$.
On the example: $p = 0.95$, $n = 10$: with checkpoints $10 / 0.95 \approx 10.5$ step runs; restarting from scratch $(1 - 0.599) / (0.05 \times 0.599) \approx 13.4$. At $n = 50$ it's about 52.6 vs. 240.
Level 3: in Python
In Python:
def with_checkpoints(n, p):
# each step needs 1/p attempts on average
return n / p
def restart_from_scratch(n, p):
return (1 - p ** n) / ((1 - p) * p ** n)
round(with_checkpoints(10, 0.95), 1), round(restart_from_scratch(10, 0.95), 1) # → (10.5, 13.4)
round(with_checkpoints(50, 0.95), 1), round(restart_from_scratch(50, 0.95)) # → (52.6, 240)
Reading it: the x-axis is task length, and the y-axis (log scale) is how many step executions you should expect to pay for. On a log scale each gridline is ten times the one below, so exponential growth draws a straight line. The restart line is that straight line: every extra step multiplies its cost by about $1/0.95 \approx 1.05$. The checkpoint line ($n/p$, just proportional to $n$) bends over and flattens, because on a log scale going from 10 to 20 steps rises no more than going from 5 to 10. The widening gap between them is the cost of having no checkpoints, and at 50 steps it's already about 240 step runs against 53. Long tasks without checkpoints aren't just unreliable; they're expensive. Durable execution is the engineering name for this: a workflow engine that saves each step's result so a crashed run resumes where it stopped.
In code: expected_step_runs evaluates both formulas: $n/p$ with
checkpoints, the restart formula without.
4. Reflection vs. external verification
Everyday picture. Proofreading your own essay versus having someone run the numbers in it. You read what you meant to write, so your own blind spots stay blind. A calculator doesn't share them.
Tiny worked example. The model writes a leap-year function:
def is_leap(year):
return year % 4 == 0 # forgets that 1900 was not a leap year
Reflection (asking a model to review it) replies "Looks correct: years
divisible by 4 are leap years." External verification (running it against
known answers) replies is_leap(1900) returned True, expected False. Fed
that failure, the model's second draft passes every case.
flowchart LR G[Model writes code] --> R1[Reflection:<br/>model reads its own code] R1 -->|same blind spot| A1[Approved, still wrong] G --> V[Verification:<br/>run tests / check schema / query the DB] V -->|1900 fails| F[Failure fed back] F --> G2[Second draft] G2 --> V2[Checks pass]
Reading it: two paths leave the same first draft. The top path stays inside the model and ends at "approved, still wrong". The bottom path goes through something outside the model (tests, a schema validator, a database query), and that something produces a concrete, checkable failure the model can fix.
In code: self_review is the top path (a model reads the code and
approves or not); run_checks is the bottom path (it runs the code against
known answers); generate_until_checks_pass loops the bottom path, feeding
each failure back to the model until the checks pass.
With verification and retries, the per-step success rate rises:
Level 3: the formula and its symbols
$$ p_{\text{step}} = p \sum_{i=0}^{r} \big((1-p)\,d\big)^{i} $$
Symbols
| Symbol | Meaning |
|---|---|
| $p$ | chance one attempt succeeds |
| $d$ | chance the check detects a failed attempt |
| $r$ | retries allowed after a detected failure |
In words: you succeed on the first try, or fail and notice and succeed on the next try, and so on. Failures you don't notice get no retry.
On the example: $p = 0.95$, perfect check $d = 1$, one retry: $0.95 + 0.05 \times 0.95 = 0.9975$ per step, so ten steps succeed $0.9975^{10} \approx 97.5\%$ of the time instead of 60%. With a check that catches only half the failures ($d = 0.5$): $0.97375^{10} \approx 77\%$.
Level 3: in Python
In Python:
def p_step(p, d, r):
# p Σ_{i=0}^{r} ((1 - p) d)^i
return p * sum(((1 - p) * d) ** i for i in range(r + 1))
# a perfect check, one retry
round(p_step(0.95, 1, 1), 4) # → 0.9975
# ten such steps
round(p_step(0.95, 1, 1) ** 10, 3) # → 0.975
# a check that catches half the failures
round(p_step(0.95, 0.5, 1), 5) # → 0.97375
round(p_step(0.95, 0.5, 1) ** 10, 2) # → 0.77
Reading it: all curves use the same 95%-reliable step. The bottom curve has no checks. The middle ones add a retry after a failure is caught, with a check that catches half or all failures. The top curve allows two retries. The gap between "half" and "all" is the lesson: retries are only as good as the check that triggers them. Invest in the check.
In code: step_success computes $p_{\text{step}}$ from $p$, $d$ and $r$;
feed its result into end_to_end_success to get the ten-step curves.
In 20 seconds
- Success compounds: $p^n$. 95% per step is 60% over ten steps.
- Plan first, execute step by step, replan from the point of failure, and keep finished work.
- Decompose into subtasks with explicit checks; retry only the failed step (checkpoints).
- Self-review shares the model's blind spots; tests, schemas and database checks don't.
- Retries help only for failures you detect, so the check matters more than the retry.
Self-test questions
Q: Each step of your agent is 95% reliable. How reliable is a 20-step task, and what do you do about it? A: $0.95^{20} \approx 36\%$. Cut the steps (higher-level tools), add external checks after key steps with a retry of just that step, checkpoint so failures don't restart the run, and put human review where a mistake is expensive. With perfect checks and one retry, per-step reliability becomes 99.75%, and 20 steps succeed about 95% of the time.
Q: Why isn't "ask the model to double-check its work" enough? A: The reviewer shares the author's blind spots, and self-critique can confidently reinforce a wrong answer. Checks outside the model (tests, schema validation, a query against the source of truth, a separate grader model with a rubric) produce concrete failures the model can act on.
Q: When is plan-and-execute better than deciding one step at a time? A: For long, multi-part tasks where staying on track matters and where showing the plan to a person before acting is valuable. Always pair it with replanning. A plan is a hypothesis, and the first failed step is evidence.
Q: What makes a subtask "good"? A: It's small, concrete and verifiable. It has a definition of done that code can check, and its output is saved so a later failure doesn't redo it.
The papers behind this lesson
- Wang et al., Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models (2023). https://arxiv.org/abs/2305.04091. It showed that asking a model to first devise a plan and then carry it out step by step reduces missed steps compared with reasoning straight through.
- Shinn et al., Reflexion: Language Agents with Verbal Reinforcement Learning (2023). https://arxiv.org/abs/2303.11366. Agents that turn feedback from the environment (failed tests, wrong answers) into written lessons and retry improve markedly, and the gains come from external signals. annotated companion
Further reading
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
- Lilian Weng, LLM Powered Autonomous Agents (planning, reflection): https://lilianweng.github.io/posts/2023-06-23-agent/
- Temporal, durable execution: https://docs.temporal.io/
- LangGraph docs (persistence and checkpoints): https://langchain-ai.github.io/langgraph/
1r""" 2# Multi-step planning: plans, checks, retries and why long tasks fail 3 4Run: `python -m primer.agents.planning` 5 6This lesson builds on the step-at-a-time loop from 7`primer.agents.agent_loop`, on tools from `primer.agents.tools`, and on 8task state kept outside the model from `primer.agents.memory`. 9 10## Level 1: The practitioner's guide 11 12**In one sentence.** Multi-step planning is how an agent gets a twenty-step 13job done when each step is only mostly reliable: write the steps down, 14check each one with something outside the model, save each result so a 15failure costs one step, and revise the plan from the point of failure 16rather than starting over. 17 18**When you need it.** The moment a task takes more than a handful of 19model calls in a row. The arithmetic is unforgiving: a step that succeeds 2095% of the time gives a ten-step task a 60% chance of finishing and a 21twenty-step task 36%; at fifty steps it is under 8%. Even a 99% step 22finishes a fifty-step task only 60% of the time. That is why a demo of a 23three-step task looks great and the same agent on a real twenty-step job 24disappoints. You don't need any of this for a task that is one or two calls 25long, or for a fixed sequence your code can run without asking a model 26what to do next. The tell: an agent that finishes the small cases and 27fails the long ones, and a bill that shows it restarting from step one 28after every stumble. 29 30**Your options.** From the simplest to the most robust; each later row 31usually keeps the earlier ones: 32 33| Option | What it does | What it guarantees | What it costs | Where it lives | 34|---|---|---|---|---| 35| One step at a time | The model decides the next action after each result | Adapts to anything it sees | Drifts on long tasks; nobody sees the intent before it acts | The agent loop | 36| A fixed workflow in code | Your code chains the calls in a known order (chaining, routing, parallel branches) | Predictable path, easy to test | Only for tasks whose steps you know in advance | Your code | 37| Plan-and-execute with replanning | The model writes the whole plan first, executes it step by step, and rewrites the rest on a failure while keeping finished work | A plan a person can read before it runs; recovery that routes around the failure | An extra planning call per replan, and a cap on replans | Your loop around the model | 38| Decomposition with checks | Each subtask has a definition of done that code can check | A failure is caught where it happens | Writing a check per subtask | Your code | 39| Checkpoints | Every subtask's output is saved; a retry reruns only the failed step | A failure costs one step, not the run | Storage for each output | Your code, or a workflow engine | 40| External verification with retries | Tests, a schema validator, a database query, or a separate grader with a rubric, fed back to the model | Concrete failures the model can fix; retries that actually help | The check itself, and one more call per retry | Outside the model | 41| Human review at critical points | A person approves before an expensive or irreversible step | Mistakes stop before they cost | Latency and attention | Your product | 42 43**How to choose.** Start with the simplest thing and add a row only when 44the numbers say so. 45 46- Steps you can list in advance: a workflow in code. Anthropic's guidance 47 is to find the simplest solution and add complexity only when needed, 48 reserving agents for open-ended problems where the steps cannot be 49 predicted. 50- Steps the model must discover: plan-and-execute, with replanning and a 51 replan limit so a planner that cannot adapt doesn't loop forever. In this 52 lesson the first plan uses a retired API, the second routes around it, 53 and the step that already succeeded is not run again. 54- Any task longer than a few steps: decompose into subtasks with checks, 55 and checkpoint. The expected work to finish a ten-step task at 95% per 56 step is 10.5 step runs with checkpoints against 13.4 restarting from 57 scratch; at fifty steps it is 53 against about 240. 58- Wherever a wrong step can hide: verify outside the model. Asked to 59 review its own leap-year function, this lesson's model approves the bug; 60 a test against 1900 catches it, and the second draft passes. 61- Wherever a mistake is expensive or irreversible: a person in the loop. 62- Whatever you pick, invest in the check before the retry. With a check 63 that catches every failure and one retry, a 95% step becomes 99.75% 64 and ten steps succeed 97.5% of the time; with a check that catches half, 65 77%; with no check, 60%. Retries are only as good as the check that 66 triggers them. 67 68**What it costs.** Planning costs one extra model call up front and one per 69replan. Checks cost engineering: a test suite, a schema, a query against 70the source of truth, sometimes a second model with a rubric. Retries cost a 71call each, but only for failures that were noticed. Checkpoints cost 72storage and a small amount of bookkeeping, and they are what stop the 73cost of a long task from growing exponentially: without them every extra 74step multiplies the expected work by about 1.05, which is a straight line 75on a log scale. Human review costs latency. Against all of this, the cost 76of doing nothing is a task that is abandoned and restarted, paying for the 77same steps over and over. 78 79**What breaks.** 80 81- **A plan reality has invalidated.** The API changed, the file moved, the 82 assumption was wrong. Treat the plan as a hypothesis and the first 83 failed step as evidence; replan from there, keeping what is done. 84- **Self-review that approves the bug.** The reviewer shares the author's 85 blind spots and can confidently reinforce a wrong answer. Put the check 86 outside the model. 87- **Retrying a failure nobody detected.** A silent failure gets no retry 88 and poisons every later step. Measure your check's detection rate; it 89 matters more than the retry count. 90- **Restarting from step one.** Without checkpoints a late failure throws 91 away all the work, and long runs become both unreliable and expensive. 92 Save every output; rerun only the failed step. 93- **A planner that loops.** Cap the replans and the total steps, and 94 report the failure instead of trying forever. 95- **Subtasks that cannot be checked.** "Reconcile the invoices" has no 96 definition of done; "each invoice appears exactly once in the match" does. 97 Make every subtask small, concrete and verifiable. 98 99**In the wild.** Wang et al. (2023), *Plan-and-Solve Prompting*, showed 100that asking a model to devise a plan and then carry it out cuts the 101missing-step errors of reasoning straight through. Shinn et al. (2023), 102*Reflexion*, turn external feedback such as failed tests into written 103lessons kept in memory for the next attempt, and report 91% pass@1 on 104HumanEval against 80% for the base model; the gains come from signals 105outside the model. Anthropic's *Building effective agents* names the 106workflow patterns in the table (prompt chaining, routing, parallelization, 107orchestrator-workers, evaluator-optimizer) and insists on stopping 108conditions such as a maximum number of iterations. Temporal is durable 109execution as a product: it persists a complete event history, and when a 110process dies another rebuilds the state and resumes where it stopped, with 111local variables intact. LangGraph's persistence and checkpoints are the 112same idea inside an agent framework. 113 114**Go deeper.** Level 2 builds the arithmetic and the remedies in plain 115Python: the compounding formula and its curves, a plan-and-execute loop 116that replans from a failed step, five subtasks with checks and the retry 117that reruns only one of them, the expected-cost formulas with and without 118checkpoints, and a reflection-versus-verification experiment on a 119leap-year function. If you only needed to choose, you are done. 120 121## Level 2: How it works, from scratch 122 123What follows measures the arithmetic, then builds the three remedies one 124mechanism at a time. 125 126An agent that does one thing is easy. An agent that does twenty things in a 127row mostly fails, for a reason that is pure arithmetic. This lesson measures 128that arithmetic, then builds the three remedies: **decompose** the task into 129steps you can check, **verify** each step with something outside the model, 130and **checkpoint** so a failure costs one step instead of the whole run. 131 132## 1. Compounding error: why long tasks fail 133 134**Everyday picture.** A relay race with ten runners. Each runner drops the 135baton only 1 time in 20 during their leg. That sounds safe, but the team 136needs *all ten* legs to go cleanly, and it loses about 4 races in 10. 137 138**Tiny worked example.** Each step of an agent succeeds 95% of the time. 139 140| steps | chance all succeed | 141|---|---| 142| 1 | 0.95 | 143| 2 | 0.95 × 0.95 = 0.9025 | 144| 10 | 0.95¹⁰ ≈ 0.599 | 145| 20 | 0.95²⁰ ≈ 0.358 | 146 147$$ 148P(\text{task succeeds}) = p^{\,n} 149$$ 150 151**Symbols** 152 153| Symbol | Meaning | 154|---|---| 155| $p$ | chance a single step succeeds | 156| $n$ | number of steps the task needs | 157 158**In words:** multiply the per-step success rate by itself once per step. 159 160**On the example:** $0.95^{10} \approx 0.60$ and $0.95^{20} \approx 0.36$. 161A 95%-reliable step is a 60%-reliable 10-step task. 162 163**In Python:** 164 165```python 166p = 0.95 167# ten steps that must all succeed 168round(p ** 10, 2) # → 0.6 169round(p ** 20, 2) # → 0.36 170``` 171 172 173 174**Reading it:** the x-axis is the number of steps in the task, and the y-axis is the 175chance the whole task succeeds. Every curve starts near the top and slides 176down, and the lower the per-step rate, the faster it falls. Find 10 steps on the 177x-axis and read up to the 95% curve: about 0.6. That's why a demo of a 178three-step task looks great and the same agent on a real twenty-step job 179disappoints. 180 181**In code:** `end_to_end_success` multiplies the per-step rate by itself once 182per step, which is the whole formula above. 183 184**Why it matters.** The remedies all attack $p$ or $n$: fewer steps 185(higher-level tools, see `primer.agents.tools`), checks that catch failures, 186retries of just the failed step, and human review at critical points. 187 188## 2. Plan-and-execute, with replanning 189 190**Everyday picture.** Cooking from a recipe. You read the whole recipe first 191(the *plan*), then do the steps. Halfway through you discover you're out of 192eggs. You don't start over, and you don't pretend you have eggs. You revise 193the rest of the recipe with what you've got, keeping everything you've 194already chopped. 195 196**Tiny worked example.** Goal "Reconcile Q3 invoices". The planner's first 197plan uses an old API. Here's the actual run: 198 199```text 200planner -> {"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]} 201execute fetch_payments ok 202execute fetch_invoices_v1 failed: invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2 203planner <- "Completed so far: fetch_payments. Step fetch_invoices_v1 failed: ...retired..." 204planner -> {"steps": ["fetch_invoices_v2", "match", "draft_summary"]} 205execute fetch_invoices_v2 ok 206execute match ok 207execute draft_summary ok (fetch_payments was NOT run again) 208``` 209 210```mermaid 211flowchart TD 212 G[Goal] --> P[Planner writes a plan] 213 P --> X[Execute next step] 214 X --> C{Step result<br/>meets expectations?} 215 C -->|yes| M{More steps?} 216 M -->|yes| X 217 M -->|no| D[Done] 218 C -->|no| R{Replans left?} 219 R -->|yes| F[Tell the planner:<br/>what's done + why it failed] 220 F --> P 221 R -->|no| S[Stop: report the failure] 222``` 223 224**Reading it:** the inner loop (execute → check → next) is the plan being 225followed. The outer loop back to *Planner* only happens on a failed check, 226and it carries two facts: what's already done (so work is kept) and why the 227step failed (so the new plan can route around it). The *Replans left?* 228diamond stops a planner that can't adapt from looping forever. 229 230**The code.** `plan_and_execute(planner, goal)` asks for a JSON plan, runs 231each action, records outputs in `state`, and on `StepFailed` asks for a 232revised plan. Steps already in `state` are skipped. 233 234**In code:** `PlanRun` is what `plan_and_execute` returns: every plan the 235planner wrote, each executed step with "ok" or its failure, and the replan count. 236 237**Why it matters.** Deciding one step at a time (a pure ReAct loop, see 238`primer.agents.agent_loop`) drifts on long tasks. An upfront plan keeps the 239agent on track and lets a person see its intent before it acts. The risk is 240following a plan that reality has invalidated, which is why replanning exists. 241 242## 3. Decomposition into verifiable subtasks 243 244**Everyday picture.** Moving house. "Move house" isn't something you can check off, 245but "every box labelled", "van booked for Saturday" and "keys handed back" 246are. Each has a clear definition of done. 247 248**Tiny worked example.** "Reconcile Q3 invoices" becomes five subtasks, each 249with a check: 250 251| subtask | output | check (definition of done) | 252|---|---|---| 253| fetch_invoices | 4 invoices | at least one invoice returned | 254| fetch_payments | 3 payments | at least one payment returned | 255| match | 4 rows: invoiced vs paid | each invoice appears exactly once | 256| list_mismatches | INV-102 (1200 vs 1100), INV-104 (450 vs 0) | none | 257| draft_summary | "Q3 reconciliation: 4 invoices, 3 payments, 2 mismatches (INV-102 short by 100, INV-104 unpaid)." | mentions the mismatches | 258 259When the payments API times out once, **only `fetch_payments` runs again**: 260attempts are `{fetch_invoices: 1, fetch_payments: 2, match: 1, ...}`. 261 262```mermaid 263flowchart LR 264 A[fetch_invoices] --> C[match] 265 B[fetch_payments] --> C 266 C --> D[list_mismatches] 267 D --> E[draft_summary] 268 B -. timeout .-> B 269``` 270 271**Reading it:** arrows show which outputs feed which step. The dotted 272self-loop on `fetch_payments` is a retry. Because every step's output is 273kept (a *checkpoint*), the retry doesn't redo `fetch_invoices`. 274 275**In code:** `Subtask` pairs a step's work with its check (the definition of 276done); `run_subtasks` runs them in order, retries only the step whose check 277failed, and returns a `RunReport` of outputs and attempts. `reconcile_q3` 278builds the five subtasks in the table. 279 280### Checkpoints: how much work a failure costs 281 282Without checkpoints, a failure anywhere means starting again from step one. 283With them, you retry just the failed step. The expected number of step 284executions to finish an $n$-step task: 285 286$$ 287\text{with checkpoints: } \frac{n}{p} 288\qquad 289\text{restart from scratch: } \frac{1 - p^{n}}{(1-p)\,p^{n}} 290$$ 291 292**Symbols** 293 294| Symbol | Meaning | 295|---|---| 296| $p$ | chance a single step attempt succeeds (failures are noticed) | 297| $n$ | number of steps | 298 299**In words:** with checkpoints each step needs on average $1/p$ attempts, so 300$n$ steps need $n/p$. Restarting from scratch needs $n$ successes *in a row*, 301and the expected wait for a run of $n$ successes grows roughly like $1/p^n$. 302 303**On the example:** $p = 0.95$, $n = 10$: with checkpoints 304$10 / 0.95 \approx 10.5$ step runs; restarting from scratch 305$(1 - 0.599) / (0.05 \times 0.599) \approx 13.4$. At $n = 50$ it's about 30652.6 vs. 240. 307 308**In Python:** 309 310```python 311def with_checkpoints(n, p): 312 # each step needs 1/p attempts on average 313 return n / p 314def restart_from_scratch(n, p): 315 return (1 - p ** n) / ((1 - p) * p ** n) 316round(with_checkpoints(10, 0.95), 1), round(restart_from_scratch(10, 0.95), 1) # → (10.5, 13.4) 317round(with_checkpoints(50, 0.95), 1), round(restart_from_scratch(50, 0.95)) # → (52.6, 240) 318``` 319 320 321 322**Reading it:** the x-axis is task length, and the y-axis (log scale) is how 323many step executions you should expect to pay for. On a log scale each 324gridline is ten times the one below, so exponential growth draws a 325*straight* line. The restart line is that straight line: every extra step 326multiplies its cost by about $1/0.95 \approx 1.05$. The checkpoint line 327($n/p$, just proportional to $n$) bends over and flattens, because on a log 328scale going from 10 to 20 steps rises no more than going from 5 to 10. The 329widening gap between them is the cost of having no checkpoints, and at 50 330steps it's already about 240 step runs against 53. Long tasks without 331checkpoints aren't just unreliable; they're expensive. **Durable execution** is the engineering name 332for this: a workflow engine that saves each step's result so a crashed run 333resumes where it stopped. 334 335**In code:** `expected_step_runs` evaluates both formulas: $n/p$ with 336checkpoints, the restart formula without. 337 338## 4. Reflection vs. external verification 339 340**Everyday picture.** Proofreading your own essay versus having someone run the 341numbers in it. You read what you *meant* to write, so your own blind spots 342stay blind. A calculator doesn't share them. 343 344**Tiny worked example.** The model writes a leap-year function: 345 346```python 347def is_leap(year): 348 return year % 4 == 0 # forgets that 1900 was not a leap year 349``` 350 351*Reflection* (asking a model to review it) replies "Looks correct: years 352divisible by 4 are leap years." *External verification* (running it against 353known answers) replies `is_leap(1900) returned True, expected False`. Fed 354that failure, the model's second draft passes every case. 355 356```mermaid 357flowchart LR 358 G[Model writes code] --> R1[Reflection:<br/>model reads its own code] 359 R1 -->|same blind spot| A1[Approved, still wrong] 360 G --> V[Verification:<br/>run tests / check schema / query the DB] 361 V -->|1900 fails| F[Failure fed back] 362 F --> G2[Second draft] 363 G2 --> V2[Checks pass] 364``` 365 366**Reading it:** two paths leave the same first draft. The top path stays inside 367the model and ends at "approved, still wrong". The bottom path goes through 368something outside the model (tests, a schema validator, a database query), 369and that something produces a concrete, checkable failure the model can fix. 370 371**In code:** `self_review` is the top path (a model reads the code and 372approves or not); `run_checks` is the bottom path (it runs the code against 373known answers); `generate_until_checks_pass` loops the bottom path, feeding 374each failure back to the model until the checks pass. 375 376With verification and retries, the per-step success rate rises: 377 378$$ 379p_{\text{step}} = p \sum_{i=0}^{r} \big((1-p)\,d\big)^{i} 380$$ 381 382**Symbols** 383 384| Symbol | Meaning | 385|---|---| 386| $p$ | chance one attempt succeeds | 387| $d$ | chance the check *detects* a failed attempt | 388| $r$ | retries allowed after a detected failure | 389 390**In words:** you succeed on the first try, or fail *and notice* and succeed 391on the next try, and so on. Failures you don't notice get no retry. 392 393**On the example:** $p = 0.95$, perfect check $d = 1$, one retry: 394$0.95 + 0.05 \times 0.95 = 0.9975$ per step, so ten steps succeed 395$0.9975^{10} \approx 97.5\%$ of the time instead of 60%. With a check that 396catches only half the failures ($d = 0.5$): $0.97375^{10} \approx 77\%$. 397 398**In Python:** 399 400```python 401def p_step(p, d, r): 402 # p Σ_{i=0}^{r} ((1 - p) d)^i 403 return p * sum(((1 - p) * d) ** i for i in range(r + 1)) 404# a perfect check, one retry 405round(p_step(0.95, 1, 1), 4) # → 0.9975 406# ten such steps 407round(p_step(0.95, 1, 1) ** 10, 3) # → 0.975 408# a check that catches half the failures 409round(p_step(0.95, 0.5, 1), 5) # → 0.97375 410round(p_step(0.95, 0.5, 1) ** 10, 2) # → 0.77 411``` 412 413 414 415**Reading it:** all curves use the same 95%-reliable step. The bottom curve has 416no checks. The middle ones add a retry after a failure is caught, with a 417check that catches half or all failures. The top curve allows two retries. 418The gap between "half" and "all" is the lesson: retries are only as good 419as the check that triggers them. Invest in the check. 420 421**In code:** `step_success` computes $p_{\text{step}}$ from $p$, $d$ and $r$; 422feed its result into `end_to_end_success` to get the ten-step curves. 423 424## In 20 seconds 425- Success compounds: $p^n$. 95% per step is 60% over ten steps. 426- Plan first, execute step by step, replan from the point of failure, and keep finished work. 427- Decompose into subtasks with explicit checks; retry only the failed step (checkpoints). 428- Self-review shares the model's blind spots; tests, schemas and database checks don't. 429- Retries help only for failures you detect, so the check matters more than the retry. 430 431## Self-test questions 432 433**Q: Each step of your agent is 95% reliable. How reliable is a 20-step task, and what do you do about it?** 434A: $0.95^{20} \approx 36\%$. Cut the steps (higher-level tools), add 435external checks after key steps with a retry of just that step, checkpoint 436so failures don't restart the run, and put human review where a mistake is 437expensive. With perfect checks and one retry, per-step reliability becomes 43899.75%, and 20 steps succeed about 95% of the time. 439 440**Q: Why isn't "ask the model to double-check its work" enough?** 441A: The reviewer shares the author's blind spots, and self-critique can 442confidently reinforce a wrong answer. Checks outside the model (tests, 443schema validation, a query against the source of truth, a separate grader 444model with a rubric) produce concrete failures the model can act on. 445 446**Q: When is plan-and-execute better than deciding one step at a time?** 447A: For long, multi-part tasks where staying on track matters and where 448showing the plan to a person before acting is valuable. Always pair it with 449replanning. A plan is a hypothesis, and the first failed step is evidence. 450 451**Q: What makes a subtask "good"?** 452A: It's small, concrete and verifiable. It has a definition of done that code 453can check, and its output is saved so a later failure doesn't redo it. 454 455## The papers behind this lesson 456 457- **Wang et al., *Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models* (2023).** 458 https://arxiv.org/abs/2305.04091. It showed that asking a model to first 459 devise a plan and then carry it out step by step reduces missed steps 460 compared with reasoning straight through. 461- **Shinn et al., *Reflexion: Language Agents with Verbal Reinforcement Learning* (2023).** 462 https://arxiv.org/abs/2303.11366. Agents that turn feedback from the 463 environment (failed tests, wrong answers) into written lessons and retry 464 improve markedly, and the gains come from *external* signals. 465 [annotated companion](../../papers/reflexion.html) 466 467## Further reading 468- Anthropic, *Building effective agents*: https://www.anthropic.com/engineering/building-effective-agents 469- Lilian Weng, *LLM Powered Autonomous Agents* (planning, reflection): https://lilianweng.github.io/posts/2023-06-23-agent/ 470- Temporal, durable execution: https://docs.temporal.io/ 471- LangGraph docs (persistence and checkpoints): https://langchain-ai.github.io/langgraph/ 472""" 473 474from __future__ import annotations 475 476import json 477from dataclasses import dataclass, field 478from typing import Any, Callable 479 480from primer.agents.llm import LLM, ScriptedLLM, last_user_text 481 482 483def end_to_end_success(p: float, n_steps: int) -> float: 484 """Chance that n independent steps, each succeeding with probability p, all succeed.""" 485 return p**n_steps 486 487 488def step_success(p: float, detect: float, retries: int) -> float: 489 """Chance one step ends up correct when failures are caught with probability `detect` 490 and each caught failure is retried, up to `retries` times. 491 492 Attempt 1 succeeds with p. It fails *and is noticed* with (1-p)·detect, which 493 buys another attempt. A failure nobody notices counts as a silent failure. 494 495 ```text 496 P = p · (1 + q + q² + ... + q^retries), where q = (1-p)·detect 497 ``` 498 """ 499 q = (1 - p) * detect 500 return p * sum(q**i for i in range(retries + 1)) 501 502 503def expected_step_runs(p: float, n_steps: int, checkpoints: bool) -> float: 504 """Expected number of step executions to get n steps done, retrying until success. 505 506 With checkpoints, each step retries on its own: n / p (a geometric wait per step). 507 Without them, any failure throws away all progress, so you need n successes 508 *in a row*: (1 - p^n) / ((1 - p) · p^n). 509 """ 510 if checkpoints: 511 return n_steps / p 512 return (1 - p**n_steps) / ((1 - p) * p**n_steps) 513 514 515# --------------------------------------------------------------------------- 516# Decomposition: small, verifiable subtasks with checks and checkpoints 517# --------------------------------------------------------------------------- 518 519 520@dataclass 521class Subtask: 522 """One step of a decomposed task. 523 524 `run(outputs)` receives the outputs of earlier steps (by name) and returns 525 this step's output. `check(output)` returns a problem description, or None 526 when the output is acceptable. The check is the step's definition of done. 527 """ 528 529 name: str 530 run: Callable[[dict[str, Any]], Any] 531 check: Callable[[Any], str | None] = lambda out: None 532 533 534@dataclass 535class RunReport: 536 status: str # "done" | "failed" 537 outputs: dict[str, Any] 538 attempts: dict[str, int] 539 failed_step: str | None = None 540 problem: str | None = None 541 542 543def run_subtasks(subtasks: list[Subtask], max_retries: int = 2) -> RunReport: 544 """Run subtasks in order, checking each one and retrying only the one that failed.""" 545 outputs: dict[str, Any] = {} 546 attempts: dict[str, int] = {} 547 for task in subtasks: 548 problem = None 549 for _ in range(1 + max_retries): 550 attempts[task.name] = attempts.get(task.name, 0) + 1 551 try: 552 out = task.run(outputs) 553 problem = task.check(out) 554 except Exception as e: # noqa: BLE001 (a crash is just another failed attempt) 555 problem = f"{type(e).__name__}: {e}" 556 if problem is None: 557 outputs[task.name] = out 558 break 559 if problem is not None: 560 return RunReport("failed", outputs, attempts, task.name, problem) 561 return RunReport("done", outputs, attempts) 562 563 564INVOICES = [ 565 {"invoice": "INV-101", "vendor": "Acme", "amount": 1500}, 566 {"invoice": "INV-102", "vendor": "Globex", "amount": 1200}, 567 {"invoice": "INV-103", "vendor": "Initech", "amount": 800}, 568 {"invoice": "INV-104", "vendor": "Umbrella", "amount": 450}, 569] 570PAYMENTS = [ 571 {"payment": "PAY-9001", "invoice": "INV-101", "amount": 1500}, 572 {"payment": "PAY-9002", "invoice": "INV-102", "amount": 1100}, 573 {"payment": "PAY-9003", "invoice": "INV-103", "amount": 800}, 574] 575 576 577def _match(outputs: dict[str, Any]) -> list[dict[str, Any]]: 578 paid: dict[str, int] = {} 579 for pay in outputs["fetch_payments"]: 580 paid[pay["invoice"]] = paid.get(pay["invoice"], 0) + pay["amount"] 581 return [ 582 {"invoice": inv["invoice"], "vendor": inv["vendor"], "invoiced": inv["amount"], "paid": paid.get(inv["invoice"], 0)} 583 for inv in outputs["fetch_invoices"] 584 ] 585 586 587def _summary(outputs: dict[str, Any]) -> str: 588 notes = [ 589 f"{m['invoice']} unpaid" if m["paid"] == 0 else f"{m['invoice']} short by {m['invoiced'] - m['paid']}" 590 for m in outputs["list_mismatches"] 591 ] 592 return ( 593 f"Q3 reconciliation: {len(outputs['fetch_invoices'])} invoices, {len(outputs['fetch_payments'])} payments, " 594 f"{len(notes)} mismatches ({', '.join(notes)})." 595 ) 596 597 598def reconcile_q3( 599 fetch_invoices: Callable[[], list[dict]] = lambda: list(INVOICES), 600 fetch_payments: Callable[[], list[dict]] = lambda: list(PAYMENTS), 601) -> list[Subtask]: 602 """"Reconcile Q3 invoices" decomposed into five steps, each with its own definition of done.""" 603 return [ 604 Subtask("fetch_invoices", lambda o: fetch_invoices(), lambda out: None if out else "no invoices returned for Q3"), 605 Subtask("fetch_payments", lambda o: fetch_payments(), lambda out: None if out else "no payments returned for Q3"), 606 # Each invoice must appear exactly once in the match, or rows were duplicated. 607 Subtask("match", _match, lambda out: None if len({m["invoice"] for m in out}) == len(out) else "duplicated invoices"), 608 Subtask("list_mismatches", lambda o: [m for m in o["match"] if m["paid"] != m["invoiced"]]), 609 Subtask("draft_summary", _summary, lambda out: None if "mismatch" in out else "summary does not report mismatches"), 610 ] 611 612 613# --------------------------------------------------------------------------- 614# Plan-and-execute with replanning 615# --------------------------------------------------------------------------- 616 617 618class StepFailed(Exception): 619 """A step's result didn't meet expectations. The message goes back to the planner.""" 620 621 622def _fetch_invoices_v1(state: dict[str, Any]) -> list[dict]: 623 raise StepFailed("invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2") 624 625 626def _need(state: dict[str, Any], *keys: str) -> None: 627 missing = [k for k in keys if k not in state] 628 if missing: 629 raise StepFailed(f"cannot run yet: missing {missing}; fetch them first") 630 631 632def _match_step(state: dict[str, Any]) -> list[dict]: 633 _need(state, "fetch_invoices_v2", "fetch_payments") 634 return _match({"fetch_invoices": state["fetch_invoices_v2"], "fetch_payments": state["fetch_payments"]}) 635 636 637def _summary_step(state: dict[str, Any]) -> str: 638 _need(state, "match") 639 mismatches = [m for m in state["match"] if m["paid"] != m["invoiced"]] 640 return _summary({"fetch_invoices": state["fetch_invoices_v2"], "fetch_payments": state["fetch_payments"], "list_mismatches": mismatches}) 641 642 643# action name -> function(state) -> output (stored in state under the action name) 644PLAN_ACTIONS: dict[str, Callable[[dict[str, Any]], Any]] = { 645 "fetch_invoices_v1": _fetch_invoices_v1, 646 "fetch_invoices_v2": lambda state: list(INVOICES), 647 "fetch_payments": lambda state: list(PAYMENTS), 648 "match": _match_step, 649 "draft_summary": _summary_step, 650} 651 652 653@dataclass 654class PlanRun: 655 status: str # "done" | "failed" 656 replans: int 657 executed: list[tuple[str, str]] # (action, "ok" or "failed: why") 658 plans: list[list[str]] 659 state: dict[str, Any] = field(default_factory=dict) 660 661 662def _ask_planner(planner: LLM, goal: str, actions: list[str], done: list[str], feedback: str) -> list[str]: 663 prompt = ( 664 f"Goal: {goal}\n" 665 f"Available actions: {', '.join(actions)}\n" 666 f"Completed so far: {', '.join(done) or 'nothing'}\n" 667 + (f"{feedback}\n" if feedback else "") 668 + 'Reply with JSON {"steps": [...]} listing the remaining actions in order.' 669 ) 670 resp = planner.complete(system="You are a planner. Output only JSON.", messages=[{"role": "user", "content": prompt}]) 671 return list(json.loads(resp.text)["steps"]) 672 673 674def plan_and_execute( 675 planner: LLM, 676 goal: str, 677 actions: dict[str, Callable[[dict[str, Any]], Any]] | None = None, 678 max_replans: int = 2, 679) -> PlanRun: 680 """Ask for a plan, execute it step by step, and replan from where it broke. 681 682 Completed steps are remembered in `state`, so a revised plan never redoes 683 finished work. The planner only sees what it needs: the goal, the 684 actions, what's done, and why the last step failed. 685 """ 686 actions = actions or PLAN_ACTIONS 687 state: dict[str, Any] = {} 688 executed: list[tuple[str, str]] = [] 689 plans: list[list[str]] = [] 690 replans, feedback = 0, "" 691 while True: 692 plan = _ask_planner(planner, goal, list(actions), [a for a in state], feedback) 693 plans.append(plan) 694 failure = None 695 for action in plan: 696 if action in state: 697 continue # already done in an earlier plan: keep the work 698 try: 699 if action not in actions: 700 raise StepFailed(f"unknown action {action!r}") 701 state[action] = actions[action](state) 702 executed.append((action, "ok")) 703 except StepFailed as e: 704 executed.append((action, f"failed: {e}")) 705 failure = f"Step {action} failed: {e}" 706 break 707 if failure is None: 708 return PlanRun("done", replans, executed, plans, state) 709 if replans >= max_replans: 710 return PlanRun("failed", replans, executed, plans, state) 711 replans += 1 712 feedback = failure 713 714 715def demo_planner() -> ScriptedLLM: 716 """A scripted planner that starts with the old API and adapts when told it's retired.""" 717 718 def policy(system, messages, tools): 719 if "retired" in last_user_text(messages): 720 return '{"steps": ["fetch_invoices_v2", "match", "draft_summary"]}' 721 return '{"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}' 722 723 return ScriptedLLM(policy) 724 725 726# --------------------------------------------------------------------------- 727# Reflection vs. external verification 728# --------------------------------------------------------------------------- 729 730BUGGY_LEAP = "def is_leap(year):\n return year % 4 == 0\n" 731FIXED_LEAP = "def is_leap(year):\n return year % 4 == 0 and (year % 100 != 0 or year % 400 == 0)\n" 732 733# (year, is it a leap year?) The century cases are the ones a quick read misses. 734LEAP_CASES = [(2024, True), (2023, False), (1900, False), (2000, True)] 735 736 737def self_review(critic: LLM, code: str) -> tuple[bool, str]: 738 """Ask a model to review code by reading it. Returns (approved, its note).""" 739 resp = critic.complete( 740 system="Review the code. Start your reply with 'Looks correct' or 'Bug:'.", 741 messages=[{"role": "user", "content": code}], 742 ) 743 return resp.text.startswith("Looks correct"), resp.text 744 745 746def run_checks(code: str) -> list[str]: 747 """External verification: actually run the code against known answers.""" 748 namespace: dict[str, Any] = {} 749 exec(code, namespace) # our own generated snippet, run in an isolated namespace 750 fn = namespace["is_leap"] 751 return [f"is_leap({y}) returned {fn(y)}, expected {want}" for y, want in LEAP_CASES if fn(y) != want] 752 753 754def confident_critic() -> ScriptedLLM: 755 """A reviewer that shares the author's blind spot, which is the usual failure of self-critique.""" 756 return ScriptedLLM(["Looks correct: years divisible by 4 are leap years."]) 757 758 759def fixing_coder() -> ScriptedLLM: 760 """Writes the buggy version first, then fixes it once shown a failing case.""" 761 762 def policy(system, messages, tools): 763 return FIXED_LEAP if "1900" in last_user_text(messages) else BUGGY_LEAP 764 765 return ScriptedLLM(policy) 766 767 768def generate_until_checks_pass(coder: LLM, task: str, max_attempts: int = 3) -> tuple[str, int]: 769 """Generate, verify externally, feed failures back, repeat. Returns (code, attempts).""" 770 messages: list[dict[str, Any]] = [{"role": "user", "content": task}] 771 code = "" 772 for attempt in range(1, max_attempts + 1): 773 code = coder.complete(system="Reply with Python code only.", messages=messages).text 774 failures = run_checks(code) 775 if not failures: 776 return code, attempt 777 messages += [ 778 {"role": "assistant", "content": code}, 779 {"role": "user", "content": "These checks failed:\n" + "\n".join(failures)}, 780 ] 781 return code, max_attempts 782 783 784# --------------------------------------------------------------------------- 785# Figures and walkthrough 786# --------------------------------------------------------------------------- 787 788 789def figures() -> dict: 790 """Plots computed from this lesson's own formulas (matplotlib imported here).""" 791 import matplotlib 792 793 matplotlib.use("Agg") 794 import matplotlib.pyplot as plt 795 796 figs = {} 797 steps = list(range(1, 31)) 798 799 fig, ax = plt.subplots(figsize=(7, 4)) 800 for p in (0.90, 0.95, 0.99): 801 ax.plot(steps, [end_to_end_success(p, n) for n in steps], "o-", ms=3, label=f"{p:.0%} per step") 802 ax.axvline(10, ls=":", color="gray") 803 ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds) = p^n", ylim=(0, 1.02), 804 title="Compounding error: reliable steps, unreliable tasks") 805 ax.legend() 806 figs["compounding"] = fig 807 808 fig, ax = plt.subplots(figsize=(7, 4)) 809 variants = [ 810 ("no checks", 0.95), 811 ("check catches half, 1 retry", step_success(0.95, 0.5, 1)), 812 ("check catches all, 1 retry", step_success(0.95, 1.0, 1)), 813 ("check catches all, 2 retries", step_success(0.95, 1.0, 2)), 814 ] 815 for label, ps in variants: 816 ax.plot(steps, [end_to_end_success(ps, n) for n in steps], label=f"{label} (step = {ps:.4f})") 817 ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds)", ylim=(0, 1.02), 818 title="Same 95% step, different verification") 819 ax.legend(fontsize=8) 820 figs["verification"] = fig 821 822 fig, ax = plt.subplots(figsize=(7, 4)) 823 ns = list(range(1, 61)) 824 ax.plot(ns, [expected_step_runs(0.95, n, True) for n in ns], label="checkpoints: retry the failed step (n/p)") 825 ax.plot(ns, [expected_step_runs(0.95, n, False) for n in ns], label="no checkpoints: restart from step 1") 826 ax.set_yscale("log") 827 ax.set(xlabel="steps in the task (n)", ylabel="expected step executions (log scale)", 828 title="What a failure costs (p = 0.95 per step)") 829 ax.legend() 830 figs["checkpoints"] = fig 831 return figs 832 833 834def demo() -> None: 835 from primer._show import banner, say, table, takeaway 836 837 banner("1. Compounding error") 838 table(["steps", "90%/step", "95%/step", "99%/step"], 839 [(n, *(end_to_end_success(p, n) for p in (0.90, 0.95, 0.99))) for n in (1, 5, 10, 20, 50)], floatfmt=".3f") 840 takeaway("A 95%-reliable step is a 60%-reliable ten-step task.") 841 842 banner("2. Plan-and-execute with replanning") 843 run = plan_and_execute(demo_planner(), "Reconcile Q3 invoices") 844 for i, plan in enumerate(run.plans): 845 print(f" plan {i + 1}: {plan}") 846 print() 847 for action, result in run.executed: 848 print(f" {action:18} {result}") 849 print() 850 say(f"status={run.status}, replans={run.replans}. fetch_payments ran once: finished work is kept.") 851 say(f"Summary: {run.state['draft_summary']}") 852 853 banner("3. Decomposition: retry only the failed step") 854 calls: list[int] = [] 855 856 def flaky_payments() -> list[dict]: 857 calls.append(1) 858 if len(calls) == 1: 859 raise TimeoutError("payments API timed out") 860 return list(PAYMENTS) 861 862 report = run_subtasks(reconcile_q3(fetch_payments=flaky_payments)) 863 table(["subtask", "attempts"], list(report.attempts.items())) 864 for m in report.outputs["list_mismatches"]: 865 print(f" mismatch: {m}") 866 print() 867 table(["steps", "step runs with checkpoints", "step runs restarting"], 868 [(n, expected_step_runs(0.95, n, True), expected_step_runs(0.95, n, False)) for n in (5, 10, 20, 50)], floatfmt=".1f") 869 870 banner("4. Reflection vs. external verification") 871 approved, note = self_review(confident_critic(), BUGGY_LEAP) 872 print(BUGGY_LEAP) 873 say(f"Self-review: approved={approved}: {note!r}") 874 say(f"Running the test cases: {run_checks(BUGGY_LEAP)}") 875 code, attempts = generate_until_checks_pass(fixing_coder(), "Write is_leap(year).") 876 say(f"Feeding the failure back: passes on attempt {attempts}.") 877 table(["verification", "per-step", "10-step task"], 878 [(label, ps, end_to_end_success(ps, 10)) for label, ps in ( 879 ("none", 0.95), ("catches half, 1 retry", step_success(0.95, 0.5, 1)), 880 ("catches all, 1 retry", step_success(0.95, 1.0, 1)))], floatfmt=".4f") 881 takeaway("Retries are only as good as the check that triggers them. Check outside the model.") 882 883 884if __name__ == "__main__": 885 demo()
484def end_to_end_success(p: float, n_steps: int) -> float: 485 """Chance that n independent steps, each succeeding with probability p, all succeed.""" 486 return p**n_steps
Chance that n independent steps, each succeeding with probability p, all succeed.
489def step_success(p: float, detect: float, retries: int) -> float: 490 """Chance one step ends up correct when failures are caught with probability `detect` 491 and each caught failure is retried, up to `retries` times. 492 493 Attempt 1 succeeds with p. It fails *and is noticed* with (1-p)·detect, which 494 buys another attempt. A failure nobody notices counts as a silent failure. 495 496 ```text 497 P = p · (1 + q + q² + ... + q^retries), where q = (1-p)·detect 498 ``` 499 """ 500 q = (1 - p) * detect 501 return p * sum(q**i for i in range(retries + 1))
Chance one step ends up correct when failures are caught with probability detect
and each caught failure is retried, up to retries times.
Attempt 1 succeeds with p. It fails and is noticed with (1-p)·detect, which buys another attempt. A failure nobody notices counts as a silent failure.
P = p · (1 + q + q² + ... + q^retries), where q = (1-p)·detect
504def expected_step_runs(p: float, n_steps: int, checkpoints: bool) -> float: 505 """Expected number of step executions to get n steps done, retrying until success. 506 507 With checkpoints, each step retries on its own: n / p (a geometric wait per step). 508 Without them, any failure throws away all progress, so you need n successes 509 *in a row*: (1 - p^n) / ((1 - p) · p^n). 510 """ 511 if checkpoints: 512 return n_steps / p 513 return (1 - p**n_steps) / ((1 - p) * p**n_steps)
Expected number of step executions to get n steps done, retrying until success.
With checkpoints, each step retries on its own: n / p (a geometric wait per step). Without them, any failure throws away all progress, so you need n successes in a row: (1 - p^n) / ((1 - p) · p^n).
521@dataclass 522class Subtask: 523 """One step of a decomposed task. 524 525 `run(outputs)` receives the outputs of earlier steps (by name) and returns 526 this step's output. `check(output)` returns a problem description, or None 527 when the output is acceptable. The check is the step's definition of done. 528 """ 529 530 name: str 531 run: Callable[[dict[str, Any]], Any] 532 check: Callable[[Any], str | None] = lambda out: None
One step of a decomposed task.
run(outputs) receives the outputs of earlier steps (by name) and returns
this step's output. check(output) returns a problem description, or None
when the output is acceptable. The check is the step's definition of done.
535@dataclass 536class RunReport: 537 status: str # "done" | "failed" 538 outputs: dict[str, Any] 539 attempts: dict[str, int] 540 failed_step: str | None = None 541 problem: str | None = None
544def run_subtasks(subtasks: list[Subtask], max_retries: int = 2) -> RunReport: 545 """Run subtasks in order, checking each one and retrying only the one that failed.""" 546 outputs: dict[str, Any] = {} 547 attempts: dict[str, int] = {} 548 for task in subtasks: 549 problem = None 550 for _ in range(1 + max_retries): 551 attempts[task.name] = attempts.get(task.name, 0) + 1 552 try: 553 out = task.run(outputs) 554 problem = task.check(out) 555 except Exception as e: # noqa: BLE001 (a crash is just another failed attempt) 556 problem = f"{type(e).__name__}: {e}" 557 if problem is None: 558 outputs[task.name] = out 559 break 560 if problem is not None: 561 return RunReport("failed", outputs, attempts, task.name, problem) 562 return RunReport("done", outputs, attempts)
Run subtasks in order, checking each one and retrying only the one that failed.
599def reconcile_q3( 600 fetch_invoices: Callable[[], list[dict]] = lambda: list(INVOICES), 601 fetch_payments: Callable[[], list[dict]] = lambda: list(PAYMENTS), 602) -> list[Subtask]: 603 """"Reconcile Q3 invoices" decomposed into five steps, each with its own definition of done.""" 604 return [ 605 Subtask("fetch_invoices", lambda o: fetch_invoices(), lambda out: None if out else "no invoices returned for Q3"), 606 Subtask("fetch_payments", lambda o: fetch_payments(), lambda out: None if out else "no payments returned for Q3"), 607 # Each invoice must appear exactly once in the match, or rows were duplicated. 608 Subtask("match", _match, lambda out: None if len({m["invoice"] for m in out}) == len(out) else "duplicated invoices"), 609 Subtask("list_mismatches", lambda o: [m for m in o["match"] if m["paid"] != m["invoiced"]]), 610 Subtask("draft_summary", _summary, lambda out: None if "mismatch" in out else "summary does not report mismatches"), 611 ]
"Reconcile Q3 invoices" decomposed into five steps, each with its own definition of done.
619class StepFailed(Exception): 620 """A step's result didn't meet expectations. The message goes back to the planner."""
A step's result didn't meet expectations. The message goes back to the planner.
654@dataclass 655class PlanRun: 656 status: str # "done" | "failed" 657 replans: int 658 executed: list[tuple[str, str]] # (action, "ok" or "failed: why") 659 plans: list[list[str]] 660 state: dict[str, Any] = field(default_factory=dict)
675def plan_and_execute( 676 planner: LLM, 677 goal: str, 678 actions: dict[str, Callable[[dict[str, Any]], Any]] | None = None, 679 max_replans: int = 2, 680) -> PlanRun: 681 """Ask for a plan, execute it step by step, and replan from where it broke. 682 683 Completed steps are remembered in `state`, so a revised plan never redoes 684 finished work. The planner only sees what it needs: the goal, the 685 actions, what's done, and why the last step failed. 686 """ 687 actions = actions or PLAN_ACTIONS 688 state: dict[str, Any] = {} 689 executed: list[tuple[str, str]] = [] 690 plans: list[list[str]] = [] 691 replans, feedback = 0, "" 692 while True: 693 plan = _ask_planner(planner, goal, list(actions), [a for a in state], feedback) 694 plans.append(plan) 695 failure = None 696 for action in plan: 697 if action in state: 698 continue # already done in an earlier plan: keep the work 699 try: 700 if action not in actions: 701 raise StepFailed(f"unknown action {action!r}") 702 state[action] = actions[action](state) 703 executed.append((action, "ok")) 704 except StepFailed as e: 705 executed.append((action, f"failed: {e}")) 706 failure = f"Step {action} failed: {e}" 707 break 708 if failure is None: 709 return PlanRun("done", replans, executed, plans, state) 710 if replans >= max_replans: 711 return PlanRun("failed", replans, executed, plans, state) 712 replans += 1 713 feedback = failure
Ask for a plan, execute it step by step, and replan from where it broke.
Completed steps are remembered in state, so a revised plan never redoes
finished work. The planner only sees what it needs: the goal, the
actions, what's done, and why the last step failed.
716def demo_planner() -> ScriptedLLM: 717 """A scripted planner that starts with the old API and adapts when told it's retired.""" 718 719 def policy(system, messages, tools): 720 if "retired" in last_user_text(messages): 721 return '{"steps": ["fetch_invoices_v2", "match", "draft_summary"]}' 722 return '{"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}' 723 724 return ScriptedLLM(policy)
A scripted planner that starts with the old API and adapts when told it's retired.
738def self_review(critic: LLM, code: str) -> tuple[bool, str]: 739 """Ask a model to review code by reading it. Returns (approved, its note).""" 740 resp = critic.complete( 741 system="Review the code. Start your reply with 'Looks correct' or 'Bug:'.", 742 messages=[{"role": "user", "content": code}], 743 ) 744 return resp.text.startswith("Looks correct"), resp.text
Ask a model to review code by reading it. Returns (approved, its note).
747def run_checks(code: str) -> list[str]: 748 """External verification: actually run the code against known answers.""" 749 namespace: dict[str, Any] = {} 750 exec(code, namespace) # our own generated snippet, run in an isolated namespace 751 fn = namespace["is_leap"] 752 return [f"is_leap({y}) returned {fn(y)}, expected {want}" for y, want in LEAP_CASES if fn(y) != want]
External verification: actually run the code against known answers.
755def confident_critic() -> ScriptedLLM: 756 """A reviewer that shares the author's blind spot, which is the usual failure of self-critique.""" 757 return ScriptedLLM(["Looks correct: years divisible by 4 are leap years."])
A reviewer that shares the author's blind spot, which is the usual failure of self-critique.
760def fixing_coder() -> ScriptedLLM: 761 """Writes the buggy version first, then fixes it once shown a failing case.""" 762 763 def policy(system, messages, tools): 764 return FIXED_LEAP if "1900" in last_user_text(messages) else BUGGY_LEAP 765 766 return ScriptedLLM(policy)
Writes the buggy version first, then fixes it once shown a failing case.
769def generate_until_checks_pass(coder: LLM, task: str, max_attempts: int = 3) -> tuple[str, int]: 770 """Generate, verify externally, feed failures back, repeat. Returns (code, attempts).""" 771 messages: list[dict[str, Any]] = [{"role": "user", "content": task}] 772 code = "" 773 for attempt in range(1, max_attempts + 1): 774 code = coder.complete(system="Reply with Python code only.", messages=messages).text 775 failures = run_checks(code) 776 if not failures: 777 return code, attempt 778 messages += [ 779 {"role": "assistant", "content": code}, 780 {"role": "user", "content": "These checks failed:\n" + "\n".join(failures)}, 781 ] 782 return code, max_attempts
Generate, verify externally, feed failures back, repeat. Returns (code, attempts).
790def figures() -> dict: 791 """Plots computed from this lesson's own formulas (matplotlib imported here).""" 792 import matplotlib 793 794 matplotlib.use("Agg") 795 import matplotlib.pyplot as plt 796 797 figs = {} 798 steps = list(range(1, 31)) 799 800 fig, ax = plt.subplots(figsize=(7, 4)) 801 for p in (0.90, 0.95, 0.99): 802 ax.plot(steps, [end_to_end_success(p, n) for n in steps], "o-", ms=3, label=f"{p:.0%} per step") 803 ax.axvline(10, ls=":", color="gray") 804 ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds) = p^n", ylim=(0, 1.02), 805 title="Compounding error: reliable steps, unreliable tasks") 806 ax.legend() 807 figs["compounding"] = fig 808 809 fig, ax = plt.subplots(figsize=(7, 4)) 810 variants = [ 811 ("no checks", 0.95), 812 ("check catches half, 1 retry", step_success(0.95, 0.5, 1)), 813 ("check catches all, 1 retry", step_success(0.95, 1.0, 1)), 814 ("check catches all, 2 retries", step_success(0.95, 1.0, 2)), 815 ] 816 for label, ps in variants: 817 ax.plot(steps, [end_to_end_success(ps, n) for n in steps], label=f"{label} (step = {ps:.4f})") 818 ax.set(xlabel="steps in the task (n)", ylabel="P(task succeeds)", ylim=(0, 1.02), 819 title="Same 95% step, different verification") 820 ax.legend(fontsize=8) 821 figs["verification"] = fig 822 823 fig, ax = plt.subplots(figsize=(7, 4)) 824 ns = list(range(1, 61)) 825 ax.plot(ns, [expected_step_runs(0.95, n, True) for n in ns], label="checkpoints: retry the failed step (n/p)") 826 ax.plot(ns, [expected_step_runs(0.95, n, False) for n in ns], label="no checkpoints: restart from step 1") 827 ax.set_yscale("log") 828 ax.set(xlabel="steps in the task (n)", ylabel="expected step executions (log scale)", 829 title="What a failure costs (p = 0.95 per step)") 830 ax.legend() 831 figs["checkpoints"] = fig 832 return figs
Plots computed from this lesson's own formulas (matplotlib imported here).
835def demo() -> None: 836 from primer._show import banner, say, table, takeaway 837 838 banner("1. Compounding error") 839 table(["steps", "90%/step", "95%/step", "99%/step"], 840 [(n, *(end_to_end_success(p, n) for p in (0.90, 0.95, 0.99))) for n in (1, 5, 10, 20, 50)], floatfmt=".3f") 841 takeaway("A 95%-reliable step is a 60%-reliable ten-step task.") 842 843 banner("2. Plan-and-execute with replanning") 844 run = plan_and_execute(demo_planner(), "Reconcile Q3 invoices") 845 for i, plan in enumerate(run.plans): 846 print(f" plan {i + 1}: {plan}") 847 print() 848 for action, result in run.executed: 849 print(f" {action:18} {result}") 850 print() 851 say(f"status={run.status}, replans={run.replans}. fetch_payments ran once: finished work is kept.") 852 say(f"Summary: {run.state['draft_summary']}") 853 854 banner("3. Decomposition: retry only the failed step") 855 calls: list[int] = [] 856 857 def flaky_payments() -> list[dict]: 858 calls.append(1) 859 if len(calls) == 1: 860 raise TimeoutError("payments API timed out") 861 return list(PAYMENTS) 862 863 report = run_subtasks(reconcile_q3(fetch_payments=flaky_payments)) 864 table(["subtask", "attempts"], list(report.attempts.items())) 865 for m in report.outputs["list_mismatches"]: 866 print(f" mismatch: {m}") 867 print() 868 table(["steps", "step runs with checkpoints", "step runs restarting"], 869 [(n, expected_step_runs(0.95, n, True), expected_step_runs(0.95, n, False)) for n in (5, 10, 20, 50)], floatfmt=".1f") 870 871 banner("4. Reflection vs. external verification") 872 approved, note = self_review(confident_critic(), BUGGY_LEAP) 873 print(BUGGY_LEAP) 874 say(f"Self-review: approved={approved}: {note!r}") 875 say(f"Running the test cases: {run_checks(BUGGY_LEAP)}") 876 code, attempts = generate_until_checks_pass(fixing_coder(), "Write is_leap(year).") 877 say(f"Feeding the failure back: passes on attempt {attempts}.") 878 table(["verification", "per-step", "10-step task"], 879 [(label, ps, end_to_end_success(ps, 10)) for label, ps in ( 880 ("none", 0.95), ("catches half, 1 retry", step_success(0.95, 0.5, 1)), 881 ("catches all, 1 retry", step_success(0.95, 1.0, 1)))], floatfmt=".4f") 882 takeaway("Retries are only as good as the check that triggers them. Check outside the model.")