primer.agents.agent_loop

The agent loop, production-grade

Run: python -m primer.agents.agent_loop

New to the notation? primer.notation explains every symbol used here from zero. This lesson builds on the message format and tool calling in primer.agents.llm and on the autonomy spectrum in primer.agents.orchestration.

Level 1: The practitioner's guide

In one sentence. The agent loop is the piece of code that calls the model, runs the tools it asks for, sends the results back and repeats until the model says it is done or a limit you set says it is done for it.

When you need it. You need a loop whenever the model has to take more than one step whose order it decides: look something up, then act on what it found, then check. One call with one tool call you run once is not a loop, and neither is a fixed pipeline where your code decides the steps (primer.agents.orchestration). You need the production loop, rather than the ten-line version, the moment real money and real side effects are attached, because the ten-line version has no way to stop. In this lesson's happy path, one question ("How many PTO days does alice have left, and what rolls over?") takes 2 model calls and 2 tool calls run in parallel, and stops because the model said end_turn. In the same lesson, a model that keeps searching stops only because a token budget tripped: 11 calls in, the running total passes 12,000 tokens (12,990) and the run ends with an estimated cost of \$0.07. The tell that you need the controls: you cannot say, for every run that ended, why it ended.

Your options. Six ways to run the loop, from the least machinery to the most:

Option What it does What it guarantees What it costs Where it lives
No loop One model call; if it asks for a tool, run it once and call once more Bounded cost: at most two calls Nothing that needs a second decision gets done Your code
The SDK's tool runner Drives the request, run, reply cycle until the model returns no tool call or max_iterations is reached The plumbing done right: results paired by id, one message per turn, type-checked inputs Per-turn controls are limited; Claude's docs point you to the manual loop for approval, custom logging or conditional execution The provider's SDK (seven languages for Claude)
The ten-line loop Call, run tools, append results, repeat while stop_reason is tool_use Full control of every turn Every production failure is yours: it can run forever, repeat itself, crash on a tool error or execute a half-written call Your code
The production loop (this lesson) The ten-line loop plus a step limit, token and dollar budgets, loop detection, errors returned as results, a hand-off tool and parallel tool execution A named stop cause for every run, one of seven A config to tune (steps, tokens, dollars, repeat threshold) and a test per control Your code
A framework's loop The loop with tracing, persistence and hand-offs built in, and its own limit (max_turns, a recursion limit) Standard patterns and observability without writing them A dependency, its assumptions and its defaults (LangGraph's recursion limit defaults to 1,000 steps) LangGraph, the OpenAI Agents SDK, the Claude Agent SDK
A hosted loop The provider runs the loop and a sandbox for the tools; you send messages and tool results No loop code, no state files, per-session containers Session runtime on top of tokens (Claude Managed Agents lists \$0.08 per session-hour) and less say over each turn The provider's servers

How to choose. Start from what can go wrong and who pays for it.

  • A prototype, a notebook, a one-off script: the tool runner. It gets the message pairing right, which is the part people get wrong first.
  • Production with side effects (writes, emails, refunds): own the loop, or use a framework whose per-turn hooks you have read. You need an approval gate, a budget in dollars, and a hand-off to a person, in exactly your shape.
  • Many independent tool calls per turn (three lookups): make sure whatever you use runs them concurrently and returns all results in one message. The lesson's latency figure shows serial time climbing with every tool while concurrent time flattens at the slowest call.
  • Long-running or scheduled agents you would rather not host: a hosted loop, if its controls cover your approval and budget needs.
  • Whatever you pick, set every limit before the first real run: steps, tokens, dollars and a repeat threshold. This lesson's defaults are 10 steps, 50,000 tokens, and a stop when the same tool is asked for with the same arguments 3 times.

What it costs. Input tokens grow with the square of the number of steps, because each call re-sends the whole history: with 500 new tokens a step, 10 steps send 27,500 input tokens rather than the 5,000 a flat per-step cost suggests, and 20 steps send 105,000, almost four times ten steps' total. At Claude Opus 5's list prices (\$5 per million input tokens and \$25 per million output, the defaults in this lesson's AgentConfig), the runaway run above cost about 7 cents before its 12,000-token budget stopped it. Latency is one round trip per model call plus the slowest tool in each turn; a serial loop adds every tool's time instead. The controls themselves cost nothing per call: loop detection is a dictionary of (tool, arguments) counts, and a budget is a comparison.

What breaks.

  • The model never says "done". Without a step limit the loop runs until the money does. Set max_steps, and treat hitting it as an outcome to monitor, not an error to hide.
  • The same call, again and again. A model retrying one search burns a call per repeat and never progresses. Count (tool, canonical arguments) and stop at a threshold.
  • A half-written tool call. A reply cut off by max_tokens can end mid-argument. Check the stop reason before running anything; a truncated call must never execute.
  • A tool throws. Crashing the loop discards the work so far; swallowing the error makes the model guess. Return the failure as an error result with a message the model can act on ("as_of must be YYYY-MM-DD"); in the lesson, the model fixes its own call on the next step.
  • A tool the model made up. An unknown name is a KeyError in a naive loop. Answer with an error result listing the real tools.
  • Results split across messages. The API pairs each request and result by id inside one user turn; splitting them breaks the pairing and teaches the model to stop making parallel calls.
  • Guessing outside its authority. A refund over the limit needs a person. Make hand-off a tool, so it ends the run with a logged reason instead of a confident wrong answer.

In the wild. ReAct (Yao et al., 2022) named the pattern the loop implements: a short piece of reasoning, an action, an observation, repeated. Claude's SDKs ship a tool runner that loops until the model returns no tool use or max_iterations is reached, and their docs send you to the manual loop when you need approval or custom logging. The OpenAI Agents SDK runs the same call, classify, run-tools cycle and raises MaxTurnsExceeded past max_turns. LangGraph bounds a graph with a recursion limit and raises GraphRecursionError when it trips. Anthropic's Building effective agents adds the operational advice: agents trade cost and latency for open-ended capability, so test them in sandboxes and put guardrails in. Every one of these products is this lesson's diamond ("done, or budget, or step limit?") with a different name on it.

Go deeper. Level 2 traces one round trip message by message, orders the seven exits the way the loop checks them and says why that order matters, fans tool calls out and back in, and derives the quadratic cost with a formula you can rerun against the budget-burner demo. If you only needed to choose, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

The loop itself is ten lines: call the model, run the tools it asks for, append the results, repeat until it stops asking. Everything else in this lesson exists because the ten-line version fails in production, and each failure gets a named control.

The idea

Everyday picture. A new assistant runs errands for you. They can't do anything themselves. They come back after every errand and say: "here's what I found; next I'd like to do X." You carry out X and hand them the result. They decide the next step, and so on, until they say "done" (or you say "that's enough, you've spent the budget"). The assistant is the model, the errands are tool calls, and you are the loop in this file.

Tiny worked example. "How many PTO days does alice have left, and what rolls over?" Here's the real transcript the happy-path demo produces:

[user]       How many PTO days does alice have left, and what rolls over?
[assistant]  I'll check the PTO policy and alice's balance at the same time.
             tool_use toolu_1 search_kb       {"query": "PTO rollover"}
             tool_use toolu_2 get_pto_balance {"employee": "alice", "as_of": "2026-09-25"}
[user]       tool_result toolu_1  "[hr-001] Paid time off: ... Unused PTO up to 5 days rolls over. ..."
             tool_result toolu_2  {"employee": "alice", "as_of": "2026-09-25", "days_left": 12}
[assistant]  Alice has 12 PTO days left. Up to 5 unused days roll over to next year [hr-001].
             (stop_reason: end_turn)

Two model calls, two tools run in parallel, one answer. ReAct ("reason + act") is the research name for this pattern of alternating a bit of reasoning ("I'll check both at once") with actions (tool calls). The loop itself is ten lines of code. Everything else in this file exists because the ten-line version fails in production.

flowchart LR G[Goal] --> T[Think<br/>choose next step] T --> C[Call a tool] C --> O[Observe result] O --> D{Done, or budget<br/>or step limit hit?} D -->|No| T D -->|Yes| F[Final answer<br/>or hand to human]

Reading it: start at Goal on the left and follow the arrows around the cycle. The model only ever does the Think box: it looks at everything so far and picks the next action. Call a tool and Observe result are your code. The diamond is the only exit, and it has three ways out: the model says it's done, a limit trips, or the agent hands the task to a person. An agent without that diamond is a loop you can't stop.

In code: run_agent is this loop, and it returns an AgentResult holding the final text, the stop cause, every step, the usage and the cost. search_kb and get_pto_balance are the two tools in the example.

One round trip, message by message

sequenceDiagram participant U as User participant L as Your loop participant M as Model participant T as Tools U->>L: task L->>M: system + messages + tool definitions M-->>L: tool_use blocks (stop_reason = tool_use) L->>L: validate, check budget, check for loops L->>T: run the calls (concurrently) T-->>L: results or errors L->>M: ONE user message with every tool_result M-->>L: final text (stop_reason = end_turn) L-->>U: answer + stop cause

Reading it: time runs top to bottom. The model's reply is a request (tool_use), never an action. Everything between that request and the next call to the model happens in your loop, which is where every control in this file lives: validation, budgets, loop detection, concurrency. Notice the model is called twice for one question. Each call re-sends the whole history, and that's where the cost comes from.

In code: each pass through run_agent appends the model's full assistant content, runs the requested tools, and appends every result in one user message; a Step records that one model call and the ToolExecutions it triggered.

The controls, in the order the loop checks them

Everyday picture. Lending your car to a new driver: a full tank but no more fuel money (budget), "come back by six" (step limit), "if you drive past the same petrol station three times, you're lost, so come home" (loop detection), and "call me if anything feels off" (hand-off to a human).

flowchart TD R[Model response] --> S1{stop_reason?} S1 -->|refusal| X1[stop: refusal] S1 -->|max_tokens| X2[stop: max_tokens<br/>don't run half-written calls] S1 -->|end_turn| X3[stop: completed] S1 -->|tool_use| B{Over token or<br/>dollar budget?} B -->|yes| X4[stop: budget_exceeded] B -->|no| H{Called handoff_to_human?} H -->|yes| X5[stop: handoff] H -->|no| L{Same tool + same args<br/>seen N times?} L -->|yes| X6[stop: loop_detected] L -->|no| E[Run tools concurrently] E --> A[Append all results<br/>in one user message] A --> N{Step limit hit?} N -->|yes| X7[stop: max_steps] N -->|no| M[Call the model again]

Reading it: every box that starts with stop: is an exit, and each of the seven is a named stop_cause in AgentResult. Read top to bottom: first trust the model's own stop reason, then protect your wallet (budget), then honor an explicit escalation, then protect against repetition, and only then spend time running tools. The order matters. A truncated response (max_tokens) is checked before tools run, because a call cut off mid-argument must never execute.

Failure What happens Control in run_agent
Model never says "done" Runs forever max_steps
Each turn re-sends the whole conversation Cost grows quadratically with steps max_total_tokens, max_cost_usd
Model repeats the same call Burns money, no progress loop detection on (tool, canonical args)
Tool raises Loop crashes, or the model never learns why errors become is_error tool results with actionable text
Unknown tool / hallucinated name KeyError is_error result listing the real tools
Several independent calls in one turn Slow if run one by one run concurrently, return all results in one user message
Output cut off (max_tokens) Half-written tool call stop with cause max_tokens
Safety refusal (refusal) Content is not an answer stop with cause refusal
Task is outside the agent's authority Agent guesses a handoff_to_human tool that ends the run cleanly

"It stopped" is not an outcome you can monitor; "it stopped because of loop_detected at step 3" is.

In code: AgentConfig holds every limit (steps, tokens, dollars, the loop threshold) and the prices AgentConfig.cost uses to turn primer.agents.llm.Usage into dollars. A tool raises ToolError to send the model an actionable error result, and HANDOFF_TOOL_DEF is the hand-off tool's definition.

Parallel tool calls: fan out, fan in

Everyday picture. Three errands in three different shops. You can do them one after another, or send three friends at once and be done when the slowest one gets back.

flowchart LR M[Assistant turn<br/>tool_use A, tool_use B, tool_use C] --> A[run A] M --> B[run B] M --> C[run C] A --> J[One user message:<br/>tool_result A, B, C<br/>matched by tool_use_id] B --> J C --> J

Reading it: the model asked for A, B and C in the same turn, before seeing any result, so they can't depend on each other and it's safe to run them at once. Wall-clock time becomes the slowest call instead of the sum. On the way back, all three results go in a single user message. The API pairs each result with its request by id, and splitting them teaches the model to stop asking for parallel calls.

In code: execute_tools runs one turn's calls on a thread pool and returns their results in the order they were asked for, turning an unknown tool name or a raised exception into an error result instead of a crash.

Why cost grows quadratically

Each call sends the entire conversation so far. If every step adds about $t$ tokens, call $k$ sends roughly $k\,t$ input tokens, so a run of $n$ steps sends

Level 3: the formula and its symbols

$$ \text{total input} \approx \sum_{k=1}^{n} k\,t = \frac{t\,n(n+1)}{2} \approx \frac{t\,n^2}{2}. $$

Symbols

Symbol Meaning
$t$ new tokens each step adds to the conversation (tool call + result)
$k$ the step number, 1, 2, 3, ...
$n$ total steps in the run

In words: step $k$ re-sends everything from the $k$ steps before it, so the total is $t$ times $1 + 2 + \dots + n$, which is about half of $n$ squared times $t$.

On the example: with $t = 500$ and $n = 10$, the total is $500 \times 55 = 27{,}500$ input tokens, not the $5{,}000$ you'd guess. Double to $n = 20$ and it's $500 \times 210 = 105{,}000$, almost 4x.

Level 3: in Python

In Python:

t = 500
def total_input(n):
    # Σ_k k·t: step k re-sends k steps' worth
    return sum(k * t for k in range(1, n + 1))
total_input(10)  # → 27500
total_input(20)  # → 105000
# almost 4x
round(total_input(20) / total_input(10), 1)  # → 3.8

Each call of a runaway agent sends more tokens than the last, and the run stops at the first call where the running total passes the 12,000-token budget

Reading it: each bar is one call to the model in the budget-burner demo (a model that keeps searching). Bars grow step by step because the history grows. The line is the running total, which is what you pay for, and it bends upward. The dashed horizontal line is the token budget. The run stops at the first step where the total crosses it, instead of carrying on.

At 10 steps re-sending the history costs 27,500 input tokens against the 5,000 a flat per-step cost suggests, and the gap keeps widening

Reading it: the x-axis is the number of steps in a run, and the y-axis is the total input tokens sent. The straight line is what people intuitively expect ("each step costs the same"). The curve is what actually happens when every call re-sends the history. At 10 steps the gap is about 5x, and it keeps widening. That gap is why step limits, trimming tool output, and prompt caching (primer.agents.cost) matter.

Run one after another, tool calls take the sum of their latencies and keep climbing; run concurrently, they take only as long as the slowest call

Reading it: the x-axis is how many tools the model requested in one turn, each with a realistic latency between 0.2 s and 1.2 s. Run one after another, the time is the sum and climbs with every tool. Run concurrently, it's the max, which flattens out near the slowest single call. Parallel tool calls are among the cheapest latency wins in agent systems.

In code: cumulative_input_tokens evaluates the formula above, $t$ times $n(n+1)/2$, for any run length.

Running it against a real model

The loop only depends on the LLM protocol, so the same code runs against Claude:

from primer.agents.llm import ClaudeLLM
from primer.agents.agent_loop import run_agent, AgentConfig

result = run_agent(
    ClaudeLLM(),                      # needs `pip install anthropic` + credentials
    task="How many PTO days does alice have left, and what's the rollover rule?",
    tools=HR_TOOLS,                   # name -> python function
    tool_defs=HR_TOOL_DEFS,           # Anthropic tool definitions (JSON Schema)
    system="You are an HR assistant. Use tools; cite doc ids.",
    config=AgentConfig(max_steps=8, max_cost_usd=0.50),
)
print(result.stop_cause, result.final_text)

The SDK also ships a tool runner (client.beta.messages.tool_runner) that drives this loop for you. Owning the loop, as here, is what you do when you need custom budgets, loop detection, approval gates or tracing in exactly your shape.

In 20 seconds

  • The agent loop is: call model, run the tools it asks for, append results, repeat until end_turn.
  • The model never executes anything. Your code validates and runs every tool call.
  • Production loops need step limits, token/cost budgets, loop detection, error-as-result handling, and a human hand-off.
  • Parallel tool calls: run them concurrently, return all results in a single user message.
  • Input cost grows roughly with the square of the number of steps, because every call re-sends the history.

Self-test questions

Q: An agent in production suddenly costs 5x more per task. What do you check first? A: Steps per task and tokens per step, from traces. A jump usually means a loop (the same call repeated), a tool that started returning huge payloads, or a prompt change that stopped the model from recognizing "done". Step budgets and loop detection cap the damage; alerts on tokens-per-task catch it early.

Q: A tool throws an exception mid-run. What should the loop do? A: Catch it and return a tool_result with is_error: true and a message the model can act on ("date must be YYYY-MM-DD, got 'next friday'"). The model usually fixes its next call. Crashing the loop throws away the work so far; swallowing the error silently makes the model guess.

Q: Why return all parallel tool results in one message? A: The API pairs each tool_use with its tool_result by id in the next user turn. Splitting results across several messages breaks that pairing and teaches the model to stop making parallel calls, which slows every run.

Q: When should an agent hand off to a human? A: When it lacks authority (refunds over a limit), lacks information after reasonable search, detects conflicting sources, or hits a budget. Make hand-off a tool, so it is an explicit, logged outcome rather than a vague final answer.

The papers behind this lesson

  • Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022). https://arxiv.org/abs/2210.03629. It showed that interleaving short reasoning traces with tool actions, then observing the results, beats reasoning alone or acting alone, and this loop is the shape of nearly every agent built since. annotated companion

Further reading

on GitHub
  1r"""
  2# The agent loop, production-grade
  3
  4Run: `python -m primer.agents.agent_loop`
  5
  6New to the notation? `primer.notation` explains every symbol used here from
  7zero. This lesson builds on the message format and tool calling in
  8`primer.agents.llm` and on the autonomy spectrum in
  9`primer.agents.orchestration`.
 10
 11## Level 1: The practitioner's guide
 12
 13**In one sentence.** The agent loop is the piece of code that calls the
 14model, runs the tools it asks for, sends the results back and repeats until
 15the model says it is done or a limit you set says it is done for it.
 16
 17**When you need it.** You need a loop whenever the model has to take more
 18than one step whose order it decides: look something up, then act on what
 19it found, then check. One call with one tool call you run once is not a
 20loop, and neither is a fixed pipeline where your code decides the steps
 21(`primer.agents.orchestration`). You need the *production* loop, rather
 22than the ten-line version, the moment real money and real side effects are
 23attached, because the ten-line version has no way to stop. In this
 24lesson's happy path, one question ("How many PTO days does alice have
 25left, and what rolls over?") takes 2 model calls and 2 tool calls run in
 26parallel, and stops because the model said `end_turn`. In the same
 27lesson, a model that keeps searching stops only because a token budget
 28tripped: 11 calls in, the running total passes 12,000 tokens (12,990) and
 29the run ends with an estimated cost of \$0.07. The tell that you need the
 30controls: you cannot say, for every run that ended, *why* it ended.
 31
 32**Your options.** Six ways to run the loop, from the least machinery to the
 33most:
 34
 35| Option | What it does | What it guarantees | What it costs | Where it lives |
 36|---|---|---|---|---|
 37| No loop | One model call; if it asks for a tool, run it once and call once more | Bounded cost: at most two calls | Nothing that needs a second decision gets done | Your code |
 38| The SDK's tool runner | Drives the request, run, reply cycle until the model returns no tool call or `max_iterations` is reached | The plumbing done right: results paired by id, one message per turn, type-checked inputs | Per-turn controls are limited; Claude's docs point you to the manual loop for approval, custom logging or conditional execution | The provider's SDK (seven languages for Claude) |
 39| The ten-line loop | Call, run tools, append results, repeat while `stop_reason` is `tool_use` | Full control of every turn | Every production failure is yours: it can run forever, repeat itself, crash on a tool error or execute a half-written call | Your code |
 40| The production loop (this lesson) | The ten-line loop plus a step limit, token and dollar budgets, loop detection, errors returned as results, a hand-off tool and parallel tool execution | A named stop cause for every run, one of seven | A config to tune (steps, tokens, dollars, repeat threshold) and a test per control | Your code |
 41| A framework's loop | The loop with tracing, persistence and hand-offs built in, and its own limit (`max_turns`, a recursion limit) | Standard patterns and observability without writing them | A dependency, its assumptions and its defaults (LangGraph's recursion limit defaults to 1,000 steps) | LangGraph, the OpenAI Agents SDK, the Claude Agent SDK |
 42| A hosted loop | The provider runs the loop and a sandbox for the tools; you send messages and tool results | No loop code, no state files, per-session containers | Session runtime on top of tokens (Claude Managed Agents lists \$0.08 per session-hour) and less say over each turn | The provider's servers |
 43
 44**How to choose.** Start from what can go wrong and who pays for it.
 45
 46- A prototype, a notebook, a one-off script: the tool runner. It gets the
 47  message pairing right, which is the part people get wrong first.
 48- Production with side effects (writes, emails, refunds): own the loop, or
 49  use a framework whose per-turn hooks you have read. You need an approval
 50  gate, a budget in dollars, and a hand-off to a person, in exactly your
 51  shape.
 52- Many independent tool calls per turn (three lookups): make sure whatever
 53  you use runs them concurrently and returns all results in one message.
 54  The lesson's latency figure shows serial time climbing with every tool
 55  while concurrent time flattens at the slowest call.
 56- Long-running or scheduled agents you would rather not host: a hosted loop,
 57  if its controls cover your approval and budget needs.
 58- Whatever you pick, set every limit before the first real run: steps,
 59  tokens, dollars and a repeat threshold. This lesson's defaults are 10
 60  steps, 50,000 tokens, and a stop when the same tool is asked for with the
 61  same arguments 3 times.
 62
 63**What it costs.** Input tokens grow with the square of the number of
 64steps, because each call re-sends the whole history: with 500 new tokens a
 65step, 10 steps send 27,500 input tokens rather than the 5,000 a flat
 66per-step cost suggests, and 20 steps send 105,000, almost four times ten
 67steps' total. At Claude Opus 5's list prices (\$5 per million input tokens
 68and \$25 per million output, the defaults in this lesson's `AgentConfig`),
 69the runaway run above cost about 7 cents before its 12,000-token budget
 70stopped it. Latency is one round trip per model call plus the slowest tool
 71in each turn; a serial loop adds every tool's time instead. The controls
 72themselves cost nothing per call: loop detection is a dictionary of (tool,
 73arguments) counts, and a budget is a comparison.
 74
 75**What breaks.**
 76
 77- **The model never says "done".** Without a step limit the loop runs
 78  until the money does. Set `max_steps`, and treat hitting it as an outcome
 79  to monitor, not an error to hide.
 80- **The same call, again and again.** A model retrying one search burns a
 81  call per repeat and never progresses. Count (tool, canonical arguments)
 82  and stop at a threshold.
 83- **A half-written tool call.** A reply cut off by `max_tokens` can end
 84  mid-argument. Check the stop reason before running anything; a truncated
 85  call must never execute.
 86- **A tool throws.** Crashing the loop discards the work so far; swallowing
 87  the error makes the model guess. Return the failure as an error result
 88  with a message the model can act on ("as_of must be YYYY-MM-DD"); in the
 89  lesson, the model fixes its own call on the next step.
 90- **A tool the model made up.** An unknown name is a `KeyError` in a naive
 91  loop. Answer with an error result listing the real tools.
 92- **Results split across messages.** The API pairs each request and result
 93  by id inside one user turn; splitting them breaks the pairing and teaches
 94  the model to stop making parallel calls.
 95- **Guessing outside its authority.** A refund over the limit needs a
 96  person. Make hand-off a tool, so it ends the run with a logged reason
 97  instead of a confident wrong answer.
 98
 99**In the wild.** ReAct (Yao et al., 2022) named the pattern the loop
100implements: a short piece of reasoning, an action, an observation, repeated.
101Claude's SDKs ship a tool runner that loops until the model returns no tool
102use or `max_iterations` is reached, and their docs send you to the manual
103loop when you need approval or custom logging. The OpenAI Agents SDK runs
104the same call, classify, run-tools cycle and raises `MaxTurnsExceeded` past
105`max_turns`. LangGraph bounds a graph with a recursion limit and raises
106`GraphRecursionError` when it trips. Anthropic's *Building effective agents*
107adds the operational advice: agents trade cost and latency for open-ended
108capability, so test them in sandboxes and put guardrails in. Every one of
109these products is this lesson's diamond ("done, or budget, or step limit?")
110with a different name on it.
111
112**Go deeper.** Level 2 traces one round trip message by message, orders the
113seven exits the way the loop checks them and says why that order matters,
114fans tool calls out and back in, and derives the quadratic cost with a
115formula you can rerun against the budget-burner demo. If you only needed to
116choose, you are done.
117
118## Level 2: How it works, from scratch
119
120The loop itself is ten lines: call the model, run the tools it asks for,
121append the results, repeat until it stops asking. Everything else in this
122lesson exists because the ten-line version fails in production, and each
123failure gets a named control.
124
125## The idea
126
127**Everyday picture.** A new assistant runs errands for you. They can't do
128anything themselves. They come back after *every* errand and say: "here's
129what I found; next I'd like to do X." You carry out X and hand them the
130result. They decide the next step, and so on, until they say "done" (or you
131say "that's enough, you've spent the budget"). The assistant is the model,
132the errands are tool calls, and *you* are the loop in this file.
133
134**Tiny worked example.** "How many PTO days does alice have left, and what
135rolls over?" Here's the real transcript the happy-path demo produces:
136
137```text
138[user]       How many PTO days does alice have left, and what rolls over?
139[assistant]  I'll check the PTO policy and alice's balance at the same time.
140             tool_use toolu_1 search_kb       {"query": "PTO rollover"}
141             tool_use toolu_2 get_pto_balance {"employee": "alice", "as_of": "2026-09-25"}
142[user]       tool_result toolu_1  "[hr-001] Paid time off: ... Unused PTO up to 5 days rolls over. ..."
143             tool_result toolu_2  {"employee": "alice", "as_of": "2026-09-25", "days_left": 12}
144[assistant]  Alice has 12 PTO days left. Up to 5 unused days roll over to next year [hr-001].
145             (stop_reason: end_turn)
146```
147
148Two model calls, two tools run in parallel, one answer. **ReAct** ("reason +
149act") is the research name for this pattern of alternating a bit of
150reasoning ("I'll check both at once") with actions (tool calls). The loop
151itself is ten lines of code. Everything else in this file exists because
152the ten-line version fails in production.
153
154```mermaid
155flowchart LR
156  G[Goal] --> T[Think<br/>choose next step]
157  T --> C[Call a tool]
158  C --> O[Observe result]
159  O --> D{Done, or budget<br/>or step limit hit?}
160  D -->|No| T
161  D -->|Yes| F[Final answer<br/>or hand to human]
162```
163
164**Reading it:** start at *Goal* on the left and follow the arrows around the
165cycle. The model only ever does the *Think* box: it looks at everything so
166far and picks the next action. *Call a tool* and *Observe result* are your
167code. The diamond is the only exit, and it has three ways out: the model
168says it's done, a limit trips, or the agent hands the task to a person. An
169agent without that diamond is a loop you can't stop.
170
171**In code:** `run_agent` is this loop, and it returns an `AgentResult`
172holding the final text, the stop cause, every step, the usage and the cost.
173`search_kb` and `get_pto_balance` are the two tools in the example.
174
175## One round trip, message by message
176
177```mermaid
178sequenceDiagram
179  participant U as User
180  participant L as Your loop
181  participant M as Model
182  participant T as Tools
183  U->>L: task
184  L->>M: system + messages + tool definitions
185  M-->>L: tool_use blocks (stop_reason = tool_use)
186  L->>L: validate, check budget, check for loops
187  L->>T: run the calls (concurrently)
188  T-->>L: results or errors
189  L->>M: ONE user message with every tool_result
190  M-->>L: final text (stop_reason = end_turn)
191  L-->>U: answer + stop cause
192```
193
194**Reading it:** time runs top to bottom. The model's reply is a *request*
195(`tool_use`), never an action. Everything between that request and the next
196call to the model happens in your loop, which is where every control in
197this file lives: validation, budgets, loop detection, concurrency. Notice
198the model is called twice for one question. Each call re-sends the whole
199history, and that's where the cost comes from.
200
201**In code:** each pass through `run_agent` appends the model's full
202assistant content, runs the requested tools, and appends every result in one
203user message; a `Step` records that one model call and the `ToolExecution`s
204it triggered.
205
206## The controls, in the order the loop checks them
207
208**Everyday picture.** Lending your car to a new driver: a full tank but no
209more fuel money (budget), "come back by six" (step limit), "if you drive past
210the same petrol station three times, you're lost, so come home" (loop
211detection), and "call me if anything feels off" (hand-off to a human).
212
213```mermaid
214flowchart TD
215  R[Model response] --> S1{stop_reason?}
216  S1 -->|refusal| X1[stop: refusal]
217  S1 -->|max_tokens| X2[stop: max_tokens<br/>don't run half-written calls]
218  S1 -->|end_turn| X3[stop: completed]
219  S1 -->|tool_use| B{Over token or<br/>dollar budget?}
220  B -->|yes| X4[stop: budget_exceeded]
221  B -->|no| H{Called handoff_to_human?}
222  H -->|yes| X5[stop: handoff]
223  H -->|no| L{Same tool + same args<br/>seen N times?}
224  L -->|yes| X6[stop: loop_detected]
225  L -->|no| E[Run tools concurrently]
226  E --> A[Append all results<br/>in one user message]
227  A --> N{Step limit hit?}
228  N -->|yes| X7[stop: max_steps]
229  N -->|no| M[Call the model again]
230```
231
232**Reading it:** every box that starts with `stop:` is an exit, and each of
233the seven is a named `stop_cause` in `AgentResult`.
234Read top to bottom: first trust the model's own stop reason, then protect
235your wallet (budget), then honor an explicit escalation, then protect
236against repetition, and only then spend time running tools. The order
237matters. A truncated response (`max_tokens`) is checked before tools run,
238because a call cut off mid-argument must never execute.
239
240| Failure | What happens | Control in `run_agent` |
241|---|---|---|
242| Model never says "done" | Runs forever | `max_steps` |
243| Each turn re-sends the whole conversation | Cost grows quadratically with steps | `max_total_tokens`, `max_cost_usd` |
244| Model repeats the same call | Burns money, no progress | loop detection on (tool, canonical args) |
245| Tool raises | Loop crashes, or the model never learns why | errors become `is_error` tool results with actionable text |
246| Unknown tool / hallucinated name | `KeyError` | `is_error` result listing the real tools |
247| Several independent calls in one turn | Slow if run one by one | run concurrently, return **all** results in **one** user message |
248| Output cut off (`max_tokens`) | Half-written tool call | stop with cause `max_tokens` |
249| Safety refusal (`refusal`) | Content is not an answer | stop with cause `refusal` |
250| Task is outside the agent's authority | Agent guesses | a `handoff_to_human` tool that ends the run cleanly |
251
252"It stopped" is not an outcome you can monitor; "it stopped because of
253loop_detected at step 3" is.
254
255**In code:** `AgentConfig` holds every limit (steps, tokens, dollars, the
256loop threshold) and the prices `AgentConfig.cost` uses to turn `primer.agents.llm.Usage` into
257dollars. A tool raises `ToolError` to send the model an actionable error
258result, and `HANDOFF_TOOL_DEF` is the hand-off tool's definition.
259
260## Parallel tool calls: fan out, fan in
261
262**Everyday picture.** Three errands in three different shops. You can do them
263one after another, or send three friends at once and be done when the
264slowest one gets back.
265
266```mermaid
267flowchart LR
268  M[Assistant turn<br/>tool_use A, tool_use B, tool_use C] --> A[run A]
269  M --> B[run B]
270  M --> C[run C]
271  A --> J[One user message:<br/>tool_result A, B, C<br/>matched by tool_use_id]
272  B --> J
273  C --> J
274```
275
276**Reading it:** the model asked for A, B and C in the same turn, before
277seeing any result, so they can't depend on each other and it's safe to run
278them at once. Wall-clock time becomes the *slowest* call instead of the
279*sum*. On the way back, all three results go in a single user message. The
280API pairs each result with its request by id, and splitting them teaches the
281model to stop asking for parallel calls.
282
283**In code:** `execute_tools` runs one turn's calls on a thread pool and
284returns their results in the order they were asked for, turning an unknown
285tool name or a raised exception into an error result instead of a crash.
286
287## Why cost grows quadratically
288
289Each call sends the *entire* conversation so far. If every step adds about
290$t$ tokens, call $k$ sends roughly $k\,t$ input tokens, so a run of $n$ steps
291sends
292
293$$
294\text{total input} \approx \sum_{k=1}^{n} k\,t = \frac{t\,n(n+1)}{2} \approx \frac{t\,n^2}{2}.
295$$
296
297**Symbols**
298
299| Symbol | Meaning |
300|---|---|
301| $t$ | new tokens each step adds to the conversation (tool call + result) |
302| $k$ | the step number, 1, 2, 3, ... |
303| $n$ | total steps in the run |
304
305**In words:** step $k$ re-sends everything from the $k$ steps before it, so
306the total is $t$ times $1 + 2 + \dots + n$, which is about half of $n$ squared times $t$.
307
308**On the example:** with $t = 500$ and $n = 10$, the total is
309$500 \times 55 = 27{,}500$ input tokens, not the $5{,}000$ you'd guess. Double
310to $n = 20$ and it's $500 \times 210 = 105{,}000$, almost 4x.
311
312**In Python:**
313
314```python
315t = 500
316def total_input(n):
317    # Σ_k k·t: step k re-sends k steps' worth
318    return sum(k * t for k in range(1, n + 1))
319total_input(10)  # → 27500
320total_input(20)  # → 105000
321# almost 4x
322round(total_input(20) / total_input(10), 1)  # → 3.8
323```
324
325![Each call of a runaway agent sends more tokens than the last, and the run stops at the first call where the running total passes the 12,000-token budget](figures/primer.agents.agent_loop.tokens_per_step.svg)
326
327**Reading it:** each bar is one call to the model in the budget-burner demo
328(a model that keeps searching). Bars grow step by step because the history
329grows. The line is the running total, which is what you pay for, and it
330bends upward. The dashed horizontal line is the token budget. The run stops
331at the first step where the total crosses it, instead of carrying on.
332
333![At 10 steps re-sending the history costs 27,500 input tokens against the 5,000 a flat per-step cost suggests, and the gap keeps widening](figures/primer.agents.agent_loop.quadratic_cost.svg)
334
335**Reading it:** the x-axis is the number of steps in a run, and the y-axis is
336the total input tokens sent. The straight line is what people intuitively expect
337("each step costs the same"). The curve is what actually happens when every
338call re-sends the history. At 10 steps the gap is about 5x, and it keeps
339widening. That gap is why step limits, trimming tool output, and prompt
340caching (`primer.agents.cost`) matter.
341
342![Run one after another, tool calls take the sum of their latencies and keep climbing; run concurrently, they take only as long as the slowest call](figures/primer.agents.agent_loop.parallel_latency.svg)
343
344**Reading it:** the x-axis is how many tools the model requested in one turn,
345each with a realistic latency between 0.2 s and 1.2 s. Run one after another,
346the time is the *sum* and climbs with every tool. Run concurrently, it's
347the *max*, which flattens out near the slowest single call. Parallel tool
348calls are among the cheapest latency wins in agent systems.
349
350**In code:** `cumulative_input_tokens` evaluates the formula above, $t$
351times $n(n+1)/2$, for any run length.
352
353## Running it against a real model
354
355The loop only depends on the `LLM` protocol, so the same code runs against
356Claude:
357
358```python
359from primer.agents.llm import ClaudeLLM
360from primer.agents.agent_loop import run_agent, AgentConfig
361
362result = run_agent(
363    ClaudeLLM(),                      # needs `pip install anthropic` + credentials
364    task="How many PTO days does alice have left, and what's the rollover rule?",
365    tools=HR_TOOLS,                   # name -> python function
366    tool_defs=HR_TOOL_DEFS,           # Anthropic tool definitions (JSON Schema)
367    system="You are an HR assistant. Use tools; cite doc ids.",
368    config=AgentConfig(max_steps=8, max_cost_usd=0.50),
369)
370print(result.stop_cause, result.final_text)
371```
372
373The SDK also ships a *tool runner* (`client.beta.messages.tool_runner`)
374that drives this loop for you. Owning the loop, as here, is what you do
375when you need custom budgets, loop detection, approval gates or tracing in
376exactly your shape.
377
378## In 20 seconds
379- The agent loop is: call model, run the tools it asks for, append results, repeat until `end_turn`.
380- The model never executes anything. Your code validates and runs every tool call.
381- Production loops need step limits, token/cost budgets, loop detection, error-as-result handling, and a human hand-off.
382- Parallel tool calls: run them concurrently, return all results in a single user message.
383- Input cost grows roughly with the square of the number of steps, because every call re-sends the history.
384
385## Self-test questions
386
387**Q: An agent in production suddenly costs 5x more per task. What do you check first?**
388A: Steps per task and tokens per step, from traces. A jump usually means a loop
389(the same call repeated), a tool that started returning huge payloads, or a
390prompt change that stopped the model from recognizing "done". Step budgets and
391loop detection cap the damage; alerts on tokens-per-task catch it early.
392
393**Q: A tool throws an exception mid-run. What should the loop do?**
394A: Catch it and return a `tool_result` with `is_error: true` and a message the
395model can act on ("date must be YYYY-MM-DD, got 'next friday'"). The model
396usually fixes its next call. Crashing the loop throws away the work so far;
397swallowing the error silently makes the model guess.
398
399**Q: Why return all parallel tool results in one message?**
400A: The API pairs each `tool_use` with its `tool_result` by id in the next user
401turn. Splitting results across several messages breaks that pairing and
402teaches the model to stop making parallel calls, which slows every run.
403
404**Q: When should an agent hand off to a human?**
405A: When it lacks authority (refunds over a limit), lacks information after
406reasonable search, detects conflicting sources, or hits a budget. Make hand-off a
407tool, so it is an explicit, logged outcome rather than a vague final answer.
408
409## The papers behind this lesson
410
411- **Yao et al., *ReAct: Synergizing Reasoning and Acting in Language Models* (2022).**
412  https://arxiv.org/abs/2210.03629. It showed that interleaving short reasoning
413  traces with tool actions, then observing the results, beats reasoning alone or
414  acting alone, and this loop is the shape of nearly every agent built since.
415  [annotated companion](../../papers/react.html)
416
417## Further reading
418- Anthropic, *Building effective agents*: https://www.anthropic.com/engineering/building-effective-agents
419- Claude tool use, implementing the loop: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
420- Yao et al., *ReAct: Synergizing Reasoning and Acting in Language Models* (2022): https://arxiv.org/abs/2210.03629
421- Lilian Weng, *LLM Powered Autonomous Agents*: https://lilianweng.github.io/posts/2023-06-23-agent/
422"""
423
424from __future__ import annotations
425
426import json
427from collections import Counter
428from concurrent.futures import ThreadPoolExecutor
429from dataclasses import dataclass, field
430from typing import Any, Callable, Literal
431
432from primer._show import banner, say, table, takeaway
433from primer.agents.llm import (
434    LLM,
435    LLMResponse,
436    ScriptedLLM,
437    ToolCall,
438    Usage,
439    last_user_text,
440    tool_result_block,
441    tool_results,
442)
443
444StopCause = Literal[
445    "completed",  # model said end_turn with a final answer
446    "max_steps",  # step limit hit
447    "budget_exceeded",  # token or dollar budget hit
448    "loop_detected",  # same tool + same args repeated too often
449    "max_tokens",  # a response was truncated
450    "refusal",  # the model declined
451    "handoff",  # the model handed the task to a human
452]
453
454HANDOFF_TOOL = "handoff_to_human"
455
456# The hand-off tool definition. Add it to `tool_defs` to give the agent an
457# explicit, auditable way to say "this needs a person".
458HANDOFF_TOOL_DEF: dict[str, Any] = {
459    "name": HANDOFF_TOOL,
460    "description": (
461        "Escalate to a human when the request is outside your authority (for example refunds, "
462        "policy exceptions, legal questions), when sources conflict, or when you cannot find the "
463        "answer after searching. Do not use for questions you can answer with the other tools."
464    ),
465    "input_schema": {
466        "type": "object",
467        "properties": {"reason": {"type": "string", "description": "Why a human is needed, one sentence."}},
468        "required": ["reason"],
469    },
470}
471
472
473class ToolError(Exception):
474    """Raise from a tool to send an actionable `is_error` result back to the model.
475
476    The message is shown to the model verbatim, so write it for the model:
477    say what was wrong and what a valid call looks like.
478    """
479
480
481@dataclass
482class AgentConfig:
483    max_steps: int = 10
484    max_total_tokens: int = 50_000
485    max_cost_usd: float | None = None
486    # Stop when the same (tool, args) pair has been requested this many times.
487    loop_threshold: int = 3
488    max_parallel_tools: int = 4
489    max_tokens_per_call: int = 4096
490    # Dollar prices per million tokens. Defaults are Claude Opus 5 list prices
491    # at the time of writing; check the provider's pricing page.
492    price_in_per_mtok: float = 5.0
493    price_out_per_mtok: float = 25.0
494
495    def cost(self, usage: Usage) -> float:
496        return (usage.input_tokens * self.price_in_per_mtok + usage.output_tokens * self.price_out_per_mtok) / 1e6
497
498
499@dataclass
500class ToolExecution:
501    name: str
502    input: dict[str, Any]
503    content: str
504    is_error: bool
505
506
507@dataclass
508class Step:
509    """One model call plus the tools it triggered."""
510
511    index: int
512    text: str
513    tool_executions: list[ToolExecution]
514    usage: Usage
515    stop_reason: str
516
517
518@dataclass
519class AgentResult:
520    final_text: str
521    stop_cause: StopCause
522    steps: list[Step]
523    usage: Usage
524    cost_usd: float
525    transcript: list[dict[str, Any]] = field(repr=False)
526    detail: str = ""  # human-readable explanation of the stop cause
527
528    @property
529    def ok(self) -> bool:
530        return self.stop_cause == "completed"
531
532    @property
533    def tool_call_count(self) -> int:
534        return sum(len(s.tool_executions) for s in self.steps)
535
536
537def _signature(call: ToolCall) -> str:
538    # Canonical JSON (sorted keys) so {"a":1,"b":2} and {"b":2,"a":1} count as the same call.
539    return f"{call.name}:{json.dumps(call.input, sort_keys=True, default=str)}"
540
541
542def _execute_one(call: ToolCall, tools: dict[str, Callable[..., Any]]) -> ToolExecution:
543    """Run one tool call, turning every failure into an actionable error result."""
544    fn = tools.get(call.name)
545    if fn is None:
546        # Hallucinated tool names happen. Tell the model what exists.
547        return ToolExecution(call.name, call.input, f"Unknown tool '{call.name}'. Available tools: {sorted(tools)}.", True)
548    try:
549        out = fn(**call.input)
550        return ToolExecution(call.name, call.input, out if isinstance(out, str) else json.dumps(out, default=str), False)
551    except ToolError as e:
552        return ToolExecution(call.name, call.input, str(e), True)
553    except TypeError as e:
554        # Usually wrong/missing argument names. The message names the bad argument.
555        return ToolExecution(call.name, call.input, f"Bad arguments for {call.name}: {e}", True)
556    except Exception as e:  # noqa: BLE001
557        # Don't leak stack traces or internals into the context; give the class and message.
558        return ToolExecution(call.name, call.input, f"{call.name} failed ({type(e).__name__}): {e}", True)
559
560
561def execute_tools(calls: list[ToolCall], tools: dict[str, Callable[..., Any]], max_workers: int = 4) -> list[ToolExecution]:
562    """Run tool calls concurrently. Results come back in the same order as `calls`.
563
564    Tools in one assistant turn are independent by construction (the model
565    asked for them together, before seeing any result), so running them in
566    parallel is safe and cuts wall-clock time to the slowest call.
567    """
568    if len(calls) <= 1:
569        return [_execute_one(c, tools) for c in calls]
570    with ThreadPoolExecutor(max_workers=max_workers) as pool:
571        return list(pool.map(lambda c: _execute_one(c, tools), calls))
572
573
574def run_agent(
575    llm: LLM,
576    task: str,
577    tools: dict[str, Callable[..., Any]],
578    tool_defs: list[dict[str, Any]],
579    system: str = "You are a helpful assistant. Use the tools when needed.",
580    config: AgentConfig | None = None,
581) -> AgentResult:
582    """Run the think -> act -> observe loop until done or a control trips.
583
584    Args:
585        llm: anything implementing `LLM.complete` (ScriptedLLM or ClaudeLLM).
586        task: the user's request.
587        tools: tool name -> Python callable taking the tool input as kwargs.
588        tool_defs: Anthropic-format tool definitions sent to the model.
589        system: system prompt.
590        config: limits and prices.
591
592    Returns:
593        AgentResult with the final text, stop cause, per-step record, usage,
594        cost and the full transcript (for tracing and replay).
595    """
596    config = config or AgentConfig()
597    messages: list[dict[str, Any]] = [{"role": "user", "content": task}]
598    steps: list[Step] = []
599    usage = Usage()
600    seen: Counter[str] = Counter()
601
602    def finish(cause: StopCause, text: str = "", detail: str = "") -> AgentResult:
603        return AgentResult(text, cause, steps, usage, config.cost(usage), messages, detail)
604
605    for i in range(config.max_steps):
606        resp: LLMResponse = llm.complete(
607            system=system, messages=messages, tools=tool_defs, max_tokens=config.max_tokens_per_call
608        )
609        usage += resp.usage
610        # Always append the full assistant content (it may carry thinking blocks
611        # that must be sent back unchanged with the real API).
612        messages.append({"role": "assistant", "content": resp.assistant_content})
613        step = Step(i, resp.text, [], resp.usage, resp.stop_reason)
614        steps.append(step)
615
616        # --- Stop reasons that aren't "use a tool" -------------------------
617        if resp.stop_reason == "refusal":
618            return finish("refusal", resp.text, "model declined the request")
619        if resp.stop_reason == "max_tokens":
620            # A truncated response may contain a half-formed tool call. Don't run it.
621            return finish("max_tokens", resp.text, "response truncated; raise max_tokens or ask for less output")
622        if resp.stop_reason == "end_turn" or not resp.tool_calls:
623            return finish("completed", resp.text)
624
625        # --- Budget: checked after every call, before spending more ---------
626        total = usage.input_tokens + usage.output_tokens
627        if total > config.max_total_tokens:
628            return finish("budget_exceeded", resp.text, f"{total:,} tokens > budget {config.max_total_tokens:,}")
629        if config.max_cost_usd is not None and config.cost(usage) > config.max_cost_usd:
630            return finish("budget_exceeded", resp.text, f"${config.cost(usage):.4f} > budget ${config.max_cost_usd:.4f}")
631
632        # --- Hand-off: an explicit, logged outcome --------------------------
633        for call in resp.tool_calls:
634            if call.name == HANDOFF_TOOL:
635                return finish("handoff", call.input.get("reason", ""), "escalated to a human")
636
637        # --- Loop detection: same call requested too many times -------------
638        for call in resp.tool_calls:
639            seen[_signature(call)] += 1
640            if seen[_signature(call)] >= config.loop_threshold:
641                return finish(
642                    "loop_detected", resp.text, f"{call.name}({json.dumps(call.input)}) requested {seen[_signature(call)]} times"
643                )
644
645        # --- Act: run all requested tools concurrently ----------------------
646        executions = execute_tools(resp.tool_calls, tools, config.max_parallel_tools)
647        step.tool_executions = executions
648        # All results go back in ONE user message, matched by tool_use_id.
649        messages.append(
650            {
651                "role": "user",
652                "content": [
653                    tool_result_block(call.id, ex.content, ex.is_error) for call, ex in zip(resp.tool_calls, executions)
654                ],
655            }
656        )
657
658    return finish("max_steps", steps[-1].text if steps else "", f"hit max_steps={config.max_steps}")
659
660
661# ---------------------------------------------------------------------------
662# Demo tools and scripted "models"
663# ---------------------------------------------------------------------------
664
665PTO_BALANCES = {"alice": 12, "bob": 3}
666
667SEARCH_KB_DEF = {
668    "name": "search_kb",
669    "description": "Search the company knowledge base (IT, HR, finance policies). Returns doc ids and text.",
670    "input_schema": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]},
671}
672PTO_BALANCE_DEF = {
673    "name": "get_pto_balance",
674    "description": "Get an employee's remaining PTO days as of a date. Use for 'how many days do I have left'.",
675    "input_schema": {
676        "type": "object",
677        "properties": {
678            "employee": {"type": "string", "description": "lowercase username, e.g. 'alice'"},
679            "as_of": {"type": "string", "description": "date as YYYY-MM-DD"},
680        },
681        "required": ["employee", "as_of"],
682    },
683}
684
685
686def search_kb(query: str) -> str:
687    """Tiny keyword search over the shared corpus (see primer.agents.rag for the real thing)."""
688    from primer.common.corpus import DOCS
689    from primer.common.text import tokenize
690
691    q = set(tokenize(query))
692    scored = sorted(DOCS, key=lambda d: -len(q & set(tokenize(d.title + " " + d.text))))
693    return "\n".join(f"[{d.id}] {d.title}: {d.text}" for d in scored[:2])
694
695
696def get_pto_balance(employee: str, as_of: str) -> str:
697    import re
698
699    if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", as_of):
700        # Actionable: says what's wrong AND what a valid value looks like.
701        raise ToolError(f"as_of must be a date in YYYY-MM-DD format (e.g. 2026-10-02), got {as_of!r}.")
702    if employee not in PTO_BALANCES:
703        raise ToolError(f"Unknown employee {employee!r}. Known usernames: {sorted(PTO_BALANCES)}.")
704    return json.dumps({"employee": employee, "as_of": as_of, "days_left": PTO_BALANCES[employee]})
705
706
707DEMO_TOOLS = {"search_kb": search_kb, "get_pto_balance": get_pto_balance}
708DEMO_TOOL_DEFS = [SEARCH_KB_DEF, PTO_BALANCE_DEF, HANDOFF_TOOL_DEF]
709
710
711def happy_policy(system, messages, tools):
712    """Turn 1: two independent lookups in parallel. Turn 2: answer from the results."""
713    results = tool_results(messages)
714    if not results:
715        return (
716            "I'll check the PTO policy and alice's balance at the same time.",
717            [ToolCall("", "search_kb", {"query": "PTO rollover"}), ToolCall("", "get_pto_balance", {"employee": "alice", "as_of": "2026-09-25"})],
718        )
719    balance = json.loads(results[-1]["content"])["days_left"]
720    return f"Alice has {balance} PTO days left. Up to 5 unused days roll over to next year [hr-001]."
721
722
723def looping_policy(system, messages, tools):
724    """A confused model that keeps re-running the same search."""
725    return ToolCall("", "search_kb", {"query": "PTO rollover"})
726
727
728def budget_burner_policy(system, messages, tools):
729    """A model that keeps searching with new queries: no loop, just ever-growing context."""
730    n = len(tool_results(messages))
731    return ToolCall("", "search_kb", {"query": f"PTO policy details page {n}"})
732
733
734def recovering_policy(system, messages, tools):
735    """First call has a bad date; the model reads the actionable error and fixes it."""
736    results = tool_results(messages)
737    if not results:
738        return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "next friday"})
739    last = results[-1]
740    if last.get("is_error"):
741        return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "2026-10-02"})
742    return f"Bob will have {json.loads(last['content'])['days_left']} PTO days left on 2026-10-02."
743
744
745def handoff_policy(system, messages, tools):
746    if "refund" in last_user_text(messages).lower():
747        return ToolCall("", HANDOFF_TOOL, {"reason": "Refund requests over $500 need a finance approver."})
748    return "OK."
749
750
751def cumulative_input_tokens(n_steps: int, tokens_per_step: int) -> int:
752    """Total input tokens for an n-step run when step k re-sends k·t tokens of history."""
753    return tokens_per_step * n_steps * (n_steps + 1) // 2
754
755
756def figures() -> dict:
757    """Plots computed from this lesson's own code. matplotlib is imported here
758    so the lesson itself needs only NumPy."""
759    import matplotlib
760
761    matplotlib.use("Agg")
762    import matplotlib.pyplot as plt
763    import numpy as np
764
765    figs = {}
766
767    # 1. The budget-burner run, step by step.
768    budget = 12_000
769    run = run_agent(
770        ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS,
771        config=AgentConfig(max_steps=50, max_total_tokens=budget),
772    )
773    per_step = [s.usage.input_tokens + s.usage.output_tokens for s in run.steps]
774    fig, ax = plt.subplots(figsize=(7, 4))
775    ax.bar(range(len(per_step)), per_step, color="#7aa6d8", label="tokens sent this step")
776    ax.plot(range(len(per_step)), np.cumsum(per_step), "o-", color="#c0392b", label="running total")
777    ax.axhline(budget, ls="--", color="gray", label=f"budget ({budget:,})")
778    ax.set(xlabel="step", ylabel="tokens", title="A runaway agent: each call re-sends the whole history")
779    ax.legend()
780    figs["tokens_per_step"] = fig
781
782    # 2. Linear intuition vs. quadratic reality.
783    n = np.arange(1, 31)
784    t = 500
785    fig, ax = plt.subplots(figsize=(7, 4))
786    ax.plot(n, t * n, label="if each step cost the same (linear)")
787    ax.plot(n, [cumulative_input_tokens(int(k), t) for k in n], label="re-sending history (quadratic)")
788    ax.set(xlabel="steps in the run", ylabel="total input tokens", title=f"Cumulative input tokens ({t} new tokens per step)")
789    ax.legend()
790    figs["quadratic_cost"] = fig
791
792    # 3. Sequential vs. concurrent tool execution.
793    latencies = np.random.default_rng(0).uniform(0.2, 1.2, 8)
794    k = np.arange(1, 9)
795    fig, ax = plt.subplots(figsize=(7, 4))
796    ax.plot(k, [latencies[:i].sum() for i in k], "o-", label="sequential (sum of latencies)")
797    ax.plot(k, [latencies[:i].max() for i in k], "o-", label="concurrent (max latency)")
798    ax.set(xlabel="tool calls requested in one turn", ylabel="wall-clock seconds", title="Parallel tool calls: time is the slowest call, not the sum")
799    ax.legend()
800    figs["parallel_latency"] = fig
801    return figs
802
803
804def demo() -> None:
805    banner("1. Happy path: parallel tool calls, then an answer")
806    r = run_agent(ScriptedLLM(happy_policy), "How many PTO days does alice have left, and what rolls over?", DEMO_TOOLS, DEMO_TOOL_DEFS)
807    say(f"stop_cause={r.stop_cause}, steps={len(r.steps)}, tool calls={r.tool_call_count}")
808    say(f"Final answer: {r.final_text}")
809    say(
810        """
811        Step 0 asked for two tools in one turn. They ran concurrently and both
812        results went back in a single user message, which is what the API expects.
813        """
814    )
815
816    banner("2. A model that loops: caught by loop detection")
817    r = run_agent(ScriptedLLM(looping_policy), "What's the PTO rollover rule?", DEMO_TOOLS, DEMO_TOOL_DEFS)
818    say(f"stop_cause={r.stop_cause}: {r.detail}")
819    takeaway("Detect the same tool with the same canonical arguments; stop after N repeats instead of paying for N more.")
820
821    banner("3. A model that burns budget: caught by the token budget")
822    r = run_agent(
823        ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS,
824        config=AgentConfig(max_steps=50, max_total_tokens=12_000),
825    )
826    table(
827        ["step", "input tokens", "output tokens"],
828        [(s.index, s.usage.input_tokens, s.usage.output_tokens) for s in r.steps],
829    )
830    say(f"stop_cause={r.stop_cause}: {r.detail}. Estimated cost ${r.cost_usd:.4f}.")
831    say(
832        """
833        Input tokens per step climb steadily because every call re-sends the
834        whole conversation. Total input grows with the square of the step count.
835        """
836    )
837
838    banner("4. Recovering from a tool error")
839    r = run_agent(ScriptedLLM(recovering_policy), "How many PTO days will bob have on Oct 2?", DEMO_TOOLS, DEMO_TOOL_DEFS)
840    for s in r.steps:
841        for ex in s.tool_executions:
842            print(f"  step {s.index}: {ex.name}({ex.input}) -> {'ERROR: ' if ex.is_error else ''}{ex.content}")
843    print()
844    say(f"stop_cause={r.stop_cause}. Final: {r.final_text}")
845    takeaway("Errors are results, not crashes. An actionable message lets the model fix its own call.")
846
847    banner("5. Hand-off to a human")
848    r = run_agent(ScriptedLLM(handoff_policy), "Please refund my $900 conference ticket.", DEMO_TOOLS, DEMO_TOOL_DEFS)
849    say(f"stop_cause={r.stop_cause}. Reason: {r.final_text}")
850
851
852if __name__ == "__main__":
853    demo()
Level 3: the code, function by function.
StopCause = typing.Literal['completed', 'max_steps', 'budget_exceeded', 'loop_detected', 'max_tokens', 'refusal', 'handoff']
HANDOFF_TOOL = 'handoff_to_human'
HANDOFF_TOOL_DEF: dict[str, typing.Any] = {'name': 'handoff_to_human', 'description': 'Escalate to a human when the request is outside your authority (for example refunds, policy exceptions, legal questions), when sources conflict, or when you cannot find the answer after searching. Do not use for questions you can answer with the other tools.', 'input_schema': {'type': 'object', 'properties': {'reason': {'type': 'string', 'description': 'Why a human is needed, one sentence.'}}, 'required': ['reason']}}
class ToolError(builtins.Exception): on GitHub
474class ToolError(Exception):
475    """Raise from a tool to send an actionable `is_error` result back to the model.
476
477    The message is shown to the model verbatim, so write it for the model:
478    say what was wrong and what a valid call looks like.
479    """

Raise from a tool to send an actionable is_error result back to the model.

The message is shown to the model verbatim, so write it for the model: say what was wrong and what a valid call looks like.

@dataclass
class AgentConfig: on GitHub
482@dataclass
483class AgentConfig:
484    max_steps: int = 10
485    max_total_tokens: int = 50_000
486    max_cost_usd: float | None = None
487    # Stop when the same (tool, args) pair has been requested this many times.
488    loop_threshold: int = 3
489    max_parallel_tools: int = 4
490    max_tokens_per_call: int = 4096
491    # Dollar prices per million tokens. Defaults are Claude Opus 5 list prices
492    # at the time of writing; check the provider's pricing page.
493    price_in_per_mtok: float = 5.0
494    price_out_per_mtok: float = 25.0
495
496    def cost(self, usage: Usage) -> float:
497        return (usage.input_tokens * self.price_in_per_mtok + usage.output_tokens * self.price_out_per_mtok) / 1e6
AgentConfig( max_steps: int = 10, max_total_tokens: int = 50000, max_cost_usd: float | None = None, loop_threshold: int = 3, max_parallel_tools: int = 4, max_tokens_per_call: int = 4096, price_in_per_mtok: float = 5.0, price_out_per_mtok: float = 25.0)
max_steps: int = 10
max_total_tokens: int = 50000
max_cost_usd: float | None = None
loop_threshold: int = 3
max_tokens_per_call: int = 4096
price_in_per_mtok: float = 5.0
price_out_per_mtok: float = 25.0
def cost(self, usage: primer.agents.llm.Usage) -> float: on GitHub
496    def cost(self, usage: Usage) -> float:
497        return (usage.input_tokens * self.price_in_per_mtok + usage.output_tokens * self.price_out_per_mtok) / 1e6
@dataclass
class ToolExecution: on GitHub
500@dataclass
501class ToolExecution:
502    name: str
503    input: dict[str, Any]
504    content: str
505    is_error: bool
ToolExecution( name: str, input: dict[str, typing.Any], content: str, is_error: bool)
name: str
input: dict[str, typing.Any]
content: str
is_error: bool
@dataclass
class Step: on GitHub
508@dataclass
509class Step:
510    """One model call plus the tools it triggered."""
511
512    index: int
513    text: str
514    tool_executions: list[ToolExecution]
515    usage: Usage
516    stop_reason: str

One model call plus the tools it triggered.

Step( index: int, text: str, tool_executions: list[ToolExecution], usage: primer.agents.llm.Usage, stop_reason: str)
index: int
text: str
@dataclass
class AgentResult: on GitHub
519@dataclass
520class AgentResult:
521    final_text: str
522    stop_cause: StopCause
523    steps: list[Step]
524    usage: Usage
525    cost_usd: float
526    transcript: list[dict[str, Any]] = field(repr=False)
527    detail: str = ""  # human-readable explanation of the stop cause
528
529    @property
530    def ok(self) -> bool:
531        return self.stop_cause == "completed"
532
533    @property
534    def tool_call_count(self) -> int:
535        return sum(len(s.tool_executions) for s in self.steps)
AgentResult( final_text: str, stop_cause: Literal['completed', 'max_steps', 'budget_exceeded', 'loop_detected', 'max_tokens', 'refusal', 'handoff'], steps: list[Step], usage: primer.agents.llm.Usage, cost_usd: float, transcript: list[dict[str, typing.Any]], detail: str = '')
final_text: str
stop_cause: Literal['completed', 'max_steps', 'budget_exceeded', 'loop_detected', 'max_tokens', 'refusal', 'handoff']
steps: list[Step]
cost_usd: float
transcript: list[dict[str, typing.Any]]
detail: str = ''
ok: bool on GitHub
529    @property
530    def ok(self) -> bool:
531        return self.stop_cause == "completed"
tool_call_count: int on GitHub
533    @property
534    def tool_call_count(self) -> int:
535        return sum(len(s.tool_executions) for s in self.steps)
def execute_tools( calls: list[primer.agents.llm.ToolCall], tools: dict[str, typing.Callable[..., typing.Any]], max_workers: int = 4) -> list[ToolExecution]: on GitHub
562def execute_tools(calls: list[ToolCall], tools: dict[str, Callable[..., Any]], max_workers: int = 4) -> list[ToolExecution]:
563    """Run tool calls concurrently. Results come back in the same order as `calls`.
564
565    Tools in one assistant turn are independent by construction (the model
566    asked for them together, before seeing any result), so running them in
567    parallel is safe and cuts wall-clock time to the slowest call.
568    """
569    if len(calls) <= 1:
570        return [_execute_one(c, tools) for c in calls]
571    with ThreadPoolExecutor(max_workers=max_workers) as pool:
572        return list(pool.map(lambda c: _execute_one(c, tools), calls))

Run tool calls concurrently. Results come back in the same order as calls.

Tools in one assistant turn are independent by construction (the model asked for them together, before seeing any result), so running them in parallel is safe and cuts wall-clock time to the slowest call.

def run_agent( llm: primer.agents.llm.LLM, task: str, tools: dict[str, typing.Callable[..., typing.Any]], tool_defs: list[dict[str, typing.Any]], system: str = 'You are a helpful assistant. Use the tools when needed.', config: AgentConfig | None = None) -> AgentResult: on GitHub
575def run_agent(
576    llm: LLM,
577    task: str,
578    tools: dict[str, Callable[..., Any]],
579    tool_defs: list[dict[str, Any]],
580    system: str = "You are a helpful assistant. Use the tools when needed.",
581    config: AgentConfig | None = None,
582) -> AgentResult:
583    """Run the think -> act -> observe loop until done or a control trips.
584
585    Args:
586        llm: anything implementing `LLM.complete` (ScriptedLLM or ClaudeLLM).
587        task: the user's request.
588        tools: tool name -> Python callable taking the tool input as kwargs.
589        tool_defs: Anthropic-format tool definitions sent to the model.
590        system: system prompt.
591        config: limits and prices.
592
593    Returns:
594        AgentResult with the final text, stop cause, per-step record, usage,
595        cost and the full transcript (for tracing and replay).
596    """
597    config = config or AgentConfig()
598    messages: list[dict[str, Any]] = [{"role": "user", "content": task}]
599    steps: list[Step] = []
600    usage = Usage()
601    seen: Counter[str] = Counter()
602
603    def finish(cause: StopCause, text: str = "", detail: str = "") -> AgentResult:
604        return AgentResult(text, cause, steps, usage, config.cost(usage), messages, detail)
605
606    for i in range(config.max_steps):
607        resp: LLMResponse = llm.complete(
608            system=system, messages=messages, tools=tool_defs, max_tokens=config.max_tokens_per_call
609        )
610        usage += resp.usage
611        # Always append the full assistant content (it may carry thinking blocks
612        # that must be sent back unchanged with the real API).
613        messages.append({"role": "assistant", "content": resp.assistant_content})
614        step = Step(i, resp.text, [], resp.usage, resp.stop_reason)
615        steps.append(step)
616
617        # --- Stop reasons that aren't "use a tool" -------------------------
618        if resp.stop_reason == "refusal":
619            return finish("refusal", resp.text, "model declined the request")
620        if resp.stop_reason == "max_tokens":
621            # A truncated response may contain a half-formed tool call. Don't run it.
622            return finish("max_tokens", resp.text, "response truncated; raise max_tokens or ask for less output")
623        if resp.stop_reason == "end_turn" or not resp.tool_calls:
624            return finish("completed", resp.text)
625
626        # --- Budget: checked after every call, before spending more ---------
627        total = usage.input_tokens + usage.output_tokens
628        if total > config.max_total_tokens:
629            return finish("budget_exceeded", resp.text, f"{total:,} tokens > budget {config.max_total_tokens:,}")
630        if config.max_cost_usd is not None and config.cost(usage) > config.max_cost_usd:
631            return finish("budget_exceeded", resp.text, f"${config.cost(usage):.4f} > budget ${config.max_cost_usd:.4f}")
632
633        # --- Hand-off: an explicit, logged outcome --------------------------
634        for call in resp.tool_calls:
635            if call.name == HANDOFF_TOOL:
636                return finish("handoff", call.input.get("reason", ""), "escalated to a human")
637
638        # --- Loop detection: same call requested too many times -------------
639        for call in resp.tool_calls:
640            seen[_signature(call)] += 1
641            if seen[_signature(call)] >= config.loop_threshold:
642                return finish(
643                    "loop_detected", resp.text, f"{call.name}({json.dumps(call.input)}) requested {seen[_signature(call)]} times"
644                )
645
646        # --- Act: run all requested tools concurrently ----------------------
647        executions = execute_tools(resp.tool_calls, tools, config.max_parallel_tools)
648        step.tool_executions = executions
649        # All results go back in ONE user message, matched by tool_use_id.
650        messages.append(
651            {
652                "role": "user",
653                "content": [
654                    tool_result_block(call.id, ex.content, ex.is_error) for call, ex in zip(resp.tool_calls, executions)
655                ],
656            }
657        )
658
659    return finish("max_steps", steps[-1].text if steps else "", f"hit max_steps={config.max_steps}")

Run the think -> act -> observe loop until done or a control trips.

Arguments:

  • llm: anything implementing LLM.complete (ScriptedLLM or ClaudeLLM).
  • task: the user's request.
  • tools: tool name -> Python callable taking the tool input as kwargs.
  • tool_defs: Anthropic-format tool definitions sent to the model.
  • system: system prompt.
  • config: limits and prices.

Returns:

AgentResult with the final text, stop cause, per-step record, usage, cost and the full transcript (for tracing and replay).

PTO_BALANCES = {'alice': 12, 'bob': 3}
SEARCH_KB_DEF = {'name': 'search_kb', 'description': 'Search the company knowledge base (IT, HR, finance policies). Returns doc ids and text.', 'input_schema': {'type': 'object', 'properties': {'query': {'type': 'string'}}, 'required': ['query']}}
PTO_BALANCE_DEF = {'name': 'get_pto_balance', 'description': "Get an employee's remaining PTO days as of a date. Use for 'how many days do I have left'.", 'input_schema': {'type': 'object', 'properties': {'employee': {'type': 'string', 'description': "lowercase username, e.g. 'alice'"}, 'as_of': {'type': 'string', 'description': 'date as YYYY-MM-DD'}}, 'required': ['employee', 'as_of']}}
def search_kb(query: str) -> str: on GitHub
687def search_kb(query: str) -> str:
688    """Tiny keyword search over the shared corpus (see primer.agents.rag for the real thing)."""
689    from primer.common.corpus import DOCS
690    from primer.common.text import tokenize
691
692    q = set(tokenize(query))
693    scored = sorted(DOCS, key=lambda d: -len(q & set(tokenize(d.title + " " + d.text))))
694    return "\n".join(f"[{d.id}] {d.title}: {d.text}" for d in scored[:2])

Tiny keyword search over the shared corpus (see primer.agents.rag for the real thing).

def get_pto_balance(employee: str, as_of: str) -> str: on GitHub
697def get_pto_balance(employee: str, as_of: str) -> str:
698    import re
699
700    if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", as_of):
701        # Actionable: says what's wrong AND what a valid value looks like.
702        raise ToolError(f"as_of must be a date in YYYY-MM-DD format (e.g. 2026-10-02), got {as_of!r}.")
703    if employee not in PTO_BALANCES:
704        raise ToolError(f"Unknown employee {employee!r}. Known usernames: {sorted(PTO_BALANCES)}.")
705    return json.dumps({"employee": employee, "as_of": as_of, "days_left": PTO_BALANCES[employee]})
DEMO_TOOLS = {'search_kb': <function search_kb>, 'get_pto_balance': <function get_pto_balance>}
DEMO_TOOL_DEFS = [{'name': 'search_kb', 'description': 'Search the company knowledge base (IT, HR, finance policies). Returns doc ids and text.', 'input_schema': {'type': 'object', 'properties': {'query': {'type': 'string'}}, 'required': ['query']}}, {'name': 'get_pto_balance', 'description': "Get an employee's remaining PTO days as of a date. Use for 'how many days do I have left'.", 'input_schema': {'type': 'object', 'properties': {'employee': {'type': 'string', 'description': "lowercase username, e.g. 'alice'"}, 'as_of': {'type': 'string', 'description': 'date as YYYY-MM-DD'}}, 'required': ['employee', 'as_of']}}, {'name': 'handoff_to_human', 'description': 'Escalate to a human when the request is outside your authority (for example refunds, policy exceptions, legal questions), when sources conflict, or when you cannot find the answer after searching. Do not use for questions you can answer with the other tools.', 'input_schema': {'type': 'object', 'properties': {'reason': {'type': 'string', 'description': 'Why a human is needed, one sentence.'}}, 'required': ['reason']}}]
def happy_policy(system, messages, tools): on GitHub
712def happy_policy(system, messages, tools):
713    """Turn 1: two independent lookups in parallel. Turn 2: answer from the results."""
714    results = tool_results(messages)
715    if not results:
716        return (
717            "I'll check the PTO policy and alice's balance at the same time.",
718            [ToolCall("", "search_kb", {"query": "PTO rollover"}), ToolCall("", "get_pto_balance", {"employee": "alice", "as_of": "2026-09-25"})],
719        )
720    balance = json.loads(results[-1]["content"])["days_left"]
721    return f"Alice has {balance} PTO days left. Up to 5 unused days roll over to next year [hr-001]."

Turn 1: two independent lookups in parallel. Turn 2: answer from the results.

def looping_policy(system, messages, tools): on GitHub
724def looping_policy(system, messages, tools):
725    """A confused model that keeps re-running the same search."""
726    return ToolCall("", "search_kb", {"query": "PTO rollover"})

A confused model that keeps re-running the same search.

def budget_burner_policy(system, messages, tools): on GitHub
729def budget_burner_policy(system, messages, tools):
730    """A model that keeps searching with new queries: no loop, just ever-growing context."""
731    n = len(tool_results(messages))
732    return ToolCall("", "search_kb", {"query": f"PTO policy details page {n}"})

A model that keeps searching with new queries: no loop, just ever-growing context.

def recovering_policy(system, messages, tools): on GitHub
735def recovering_policy(system, messages, tools):
736    """First call has a bad date; the model reads the actionable error and fixes it."""
737    results = tool_results(messages)
738    if not results:
739        return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "next friday"})
740    last = results[-1]
741    if last.get("is_error"):
742        return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "2026-10-02"})
743    return f"Bob will have {json.loads(last['content'])['days_left']} PTO days left on 2026-10-02."

First call has a bad date; the model reads the actionable error and fixes it.

def handoff_policy(system, messages, tools): on GitHub
746def handoff_policy(system, messages, tools):
747    if "refund" in last_user_text(messages).lower():
748        return ToolCall("", HANDOFF_TOOL, {"reason": "Refund requests over $500 need a finance approver."})
749    return "OK."
def cumulative_input_tokens(n_steps: int, tokens_per_step: int) -> int: on GitHub
752def cumulative_input_tokens(n_steps: int, tokens_per_step: int) -> int:
753    """Total input tokens for an n-step run when step k re-sends k·t tokens of history."""
754    return tokens_per_step * n_steps * (n_steps + 1) // 2

Total input tokens for an n-step run when step k re-sends k·t tokens of history.

def figures() -> dict: on GitHub
757def figures() -> dict:
758    """Plots computed from this lesson's own code. matplotlib is imported here
759    so the lesson itself needs only NumPy."""
760    import matplotlib
761
762    matplotlib.use("Agg")
763    import matplotlib.pyplot as plt
764    import numpy as np
765
766    figs = {}
767
768    # 1. The budget-burner run, step by step.
769    budget = 12_000
770    run = run_agent(
771        ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS,
772        config=AgentConfig(max_steps=50, max_total_tokens=budget),
773    )
774    per_step = [s.usage.input_tokens + s.usage.output_tokens for s in run.steps]
775    fig, ax = plt.subplots(figsize=(7, 4))
776    ax.bar(range(len(per_step)), per_step, color="#7aa6d8", label="tokens sent this step")
777    ax.plot(range(len(per_step)), np.cumsum(per_step), "o-", color="#c0392b", label="running total")
778    ax.axhline(budget, ls="--", color="gray", label=f"budget ({budget:,})")
779    ax.set(xlabel="step", ylabel="tokens", title="A runaway agent: each call re-sends the whole history")
780    ax.legend()
781    figs["tokens_per_step"] = fig
782
783    # 2. Linear intuition vs. quadratic reality.
784    n = np.arange(1, 31)
785    t = 500
786    fig, ax = plt.subplots(figsize=(7, 4))
787    ax.plot(n, t * n, label="if each step cost the same (linear)")
788    ax.plot(n, [cumulative_input_tokens(int(k), t) for k in n], label="re-sending history (quadratic)")
789    ax.set(xlabel="steps in the run", ylabel="total input tokens", title=f"Cumulative input tokens ({t} new tokens per step)")
790    ax.legend()
791    figs["quadratic_cost"] = fig
792
793    # 3. Sequential vs. concurrent tool execution.
794    latencies = np.random.default_rng(0).uniform(0.2, 1.2, 8)
795    k = np.arange(1, 9)
796    fig, ax = plt.subplots(figsize=(7, 4))
797    ax.plot(k, [latencies[:i].sum() for i in k], "o-", label="sequential (sum of latencies)")
798    ax.plot(k, [latencies[:i].max() for i in k], "o-", label="concurrent (max latency)")
799    ax.set(xlabel="tool calls requested in one turn", ylabel="wall-clock seconds", title="Parallel tool calls: time is the slowest call, not the sum")
800    ax.legend()
801    figs["parallel_latency"] = fig
802    return figs

Plots computed from this lesson's own code. matplotlib is imported here so the lesson itself needs only NumPy.

def demo() -> None: on GitHub
805def demo() -> None:
806    banner("1. Happy path: parallel tool calls, then an answer")
807    r = run_agent(ScriptedLLM(happy_policy), "How many PTO days does alice have left, and what rolls over?", DEMO_TOOLS, DEMO_TOOL_DEFS)
808    say(f"stop_cause={r.stop_cause}, steps={len(r.steps)}, tool calls={r.tool_call_count}")
809    say(f"Final answer: {r.final_text}")
810    say(
811        """
812        Step 0 asked for two tools in one turn. They ran concurrently and both
813        results went back in a single user message, which is what the API expects.
814        """
815    )
816
817    banner("2. A model that loops: caught by loop detection")
818    r = run_agent(ScriptedLLM(looping_policy), "What's the PTO rollover rule?", DEMO_TOOLS, DEMO_TOOL_DEFS)
819    say(f"stop_cause={r.stop_cause}: {r.detail}")
820    takeaway("Detect the same tool with the same canonical arguments; stop after N repeats instead of paying for N more.")
821
822    banner("3. A model that burns budget: caught by the token budget")
823    r = run_agent(
824        ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS,
825        config=AgentConfig(max_steps=50, max_total_tokens=12_000),
826    )
827    table(
828        ["step", "input tokens", "output tokens"],
829        [(s.index, s.usage.input_tokens, s.usage.output_tokens) for s in r.steps],
830    )
831    say(f"stop_cause={r.stop_cause}: {r.detail}. Estimated cost ${r.cost_usd:.4f}.")
832    say(
833        """
834        Input tokens per step climb steadily because every call re-sends the
835        whole conversation. Total input grows with the square of the step count.
836        """
837    )
838
839    banner("4. Recovering from a tool error")
840    r = run_agent(ScriptedLLM(recovering_policy), "How many PTO days will bob have on Oct 2?", DEMO_TOOLS, DEMO_TOOL_DEFS)
841    for s in r.steps:
842        for ex in s.tool_executions:
843            print(f"  step {s.index}: {ex.name}({ex.input}) -> {'ERROR: ' if ex.is_error else ''}{ex.content}")
844    print()
845    say(f"stop_cause={r.stop_cause}. Final: {r.final_text}")
846    takeaway("Errors are results, not crashes. An actionable message lets the model fix its own call.")
847
848    banner("5. Hand-off to a human")
849    r = run_agent(ScriptedLLM(handoff_policy), "Please refund my $900 conference ticket.", DEMO_TOOLS, DEMO_TOOL_DEFS)
850    say(f"stop_cause={r.stop_cause}. Reason: {r.final_text}")