primer.agents.agent_loop
The agent loop, production-grade
Run: python -m primer.agents.agent_loop
New to the notation? primer.notation explains every symbol used here from
zero. This lesson builds on the message format and tool calling in
primer.agents.llm and on the autonomy spectrum in
primer.agents.orchestration.
Level 1: The practitioner's guide
In one sentence. The agent loop is the piece of code that calls the model, runs the tools it asks for, sends the results back and repeats until the model says it is done or a limit you set says it is done for it.
When you need it. You need a loop whenever the model has to take more
than one step whose order it decides: look something up, then act on what
it found, then check. One call with one tool call you run once is not a
loop, and neither is a fixed pipeline where your code decides the steps
(primer.agents.orchestration). You need the production loop, rather
than the ten-line version, the moment real money and real side effects are
attached, because the ten-line version has no way to stop. In this
lesson's happy path, one question ("How many PTO days does alice have
left, and what rolls over?") takes 2 model calls and 2 tool calls run in
parallel, and stops because the model said end_turn. In the same
lesson, a model that keeps searching stops only because a token budget
tripped: 11 calls in, the running total passes 12,000 tokens (12,990) and
the run ends with an estimated cost of \$0.07. The tell that you need the
controls: you cannot say, for every run that ended, why it ended.
Your options. Six ways to run the loop, from the least machinery to the most:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| No loop | One model call; if it asks for a tool, run it once and call once more | Bounded cost: at most two calls | Nothing that needs a second decision gets done | Your code |
| The SDK's tool runner | Drives the request, run, reply cycle until the model returns no tool call or max_iterations is reached |
The plumbing done right: results paired by id, one message per turn, type-checked inputs | Per-turn controls are limited; Claude's docs point you to the manual loop for approval, custom logging or conditional execution | The provider's SDK (seven languages for Claude) |
| The ten-line loop | Call, run tools, append results, repeat while stop_reason is tool_use |
Full control of every turn | Every production failure is yours: it can run forever, repeat itself, crash on a tool error or execute a half-written call | Your code |
| The production loop (this lesson) | The ten-line loop plus a step limit, token and dollar budgets, loop detection, errors returned as results, a hand-off tool and parallel tool execution | A named stop cause for every run, one of seven | A config to tune (steps, tokens, dollars, repeat threshold) and a test per control | Your code |
| A framework's loop | The loop with tracing, persistence and hand-offs built in, and its own limit (max_turns, a recursion limit) |
Standard patterns and observability without writing them | A dependency, its assumptions and its defaults (LangGraph's recursion limit defaults to 1,000 steps) | LangGraph, the OpenAI Agents SDK, the Claude Agent SDK |
| A hosted loop | The provider runs the loop and a sandbox for the tools; you send messages and tool results | No loop code, no state files, per-session containers | Session runtime on top of tokens (Claude Managed Agents lists \$0.08 per session-hour) and less say over each turn | The provider's servers |
How to choose. Start from what can go wrong and who pays for it.
- A prototype, a notebook, a one-off script: the tool runner. It gets the message pairing right, which is the part people get wrong first.
- Production with side effects (writes, emails, refunds): own the loop, or use a framework whose per-turn hooks you have read. You need an approval gate, a budget in dollars, and a hand-off to a person, in exactly your shape.
- Many independent tool calls per turn (three lookups): make sure whatever you use runs them concurrently and returns all results in one message. The lesson's latency figure shows serial time climbing with every tool while concurrent time flattens at the slowest call.
- Long-running or scheduled agents you would rather not host: a hosted loop, if its controls cover your approval and budget needs.
- Whatever you pick, set every limit before the first real run: steps, tokens, dollars and a repeat threshold. This lesson's defaults are 10 steps, 50,000 tokens, and a stop when the same tool is asked for with the same arguments 3 times.
What it costs. Input tokens grow with the square of the number of
steps, because each call re-sends the whole history: with 500 new tokens a
step, 10 steps send 27,500 input tokens rather than the 5,000 a flat
per-step cost suggests, and 20 steps send 105,000, almost four times ten
steps' total. At Claude Opus 5's list prices (\$5 per million input tokens
and \$25 per million output, the defaults in this lesson's AgentConfig),
the runaway run above cost about 7 cents before its 12,000-token budget
stopped it. Latency is one round trip per model call plus the slowest tool
in each turn; a serial loop adds every tool's time instead. The controls
themselves cost nothing per call: loop detection is a dictionary of (tool,
arguments) counts, and a budget is a comparison.
What breaks.
- The model never says "done". Without a step limit the loop runs
until the money does. Set
max_steps, and treat hitting it as an outcome to monitor, not an error to hide. - The same call, again and again. A model retrying one search burns a call per repeat and never progresses. Count (tool, canonical arguments) and stop at a threshold.
- A half-written tool call. A reply cut off by
max_tokenscan end mid-argument. Check the stop reason before running anything; a truncated call must never execute. - A tool throws. Crashing the loop discards the work so far; swallowing the error makes the model guess. Return the failure as an error result with a message the model can act on ("as_of must be YYYY-MM-DD"); in the lesson, the model fixes its own call on the next step.
- A tool the model made up. An unknown name is a
KeyErrorin a naive loop. Answer with an error result listing the real tools. - Results split across messages. The API pairs each request and result by id inside one user turn; splitting them breaks the pairing and teaches the model to stop making parallel calls.
- Guessing outside its authority. A refund over the limit needs a person. Make hand-off a tool, so it ends the run with a logged reason instead of a confident wrong answer.
In the wild. ReAct (Yao et al., 2022) named the pattern the loop
implements: a short piece of reasoning, an action, an observation, repeated.
Claude's SDKs ship a tool runner that loops until the model returns no tool
use or max_iterations is reached, and their docs send you to the manual
loop when you need approval or custom logging. The OpenAI Agents SDK runs
the same call, classify, run-tools cycle and raises MaxTurnsExceeded past
max_turns. LangGraph bounds a graph with a recursion limit and raises
GraphRecursionError when it trips. Anthropic's Building effective agents
adds the operational advice: agents trade cost and latency for open-ended
capability, so test them in sandboxes and put guardrails in. Every one of
these products is this lesson's diamond ("done, or budget, or step limit?")
with a different name on it.
Go deeper. Level 2 traces one round trip message by message, orders the seven exits the way the loop checks them and says why that order matters, fans tool calls out and back in, and derives the quadratic cost with a formula you can rerun against the budget-burner demo. If you only needed to choose, you are done.
Level 2: How it works, from scratch
The loop itself is ten lines: call the model, run the tools it asks for, append the results, repeat until it stops asking. Everything else in this lesson exists because the ten-line version fails in production, and each failure gets a named control.
The idea
Everyday picture. A new assistant runs errands for you. They can't do anything themselves. They come back after every errand and say: "here's what I found; next I'd like to do X." You carry out X and hand them the result. They decide the next step, and so on, until they say "done" (or you say "that's enough, you've spent the budget"). The assistant is the model, the errands are tool calls, and you are the loop in this file.
Tiny worked example. "How many PTO days does alice have left, and what rolls over?" Here's the real transcript the happy-path demo produces:
[user] How many PTO days does alice have left, and what rolls over?
[assistant] I'll check the PTO policy and alice's balance at the same time.
tool_use toolu_1 search_kb {"query": "PTO rollover"}
tool_use toolu_2 get_pto_balance {"employee": "alice", "as_of": "2026-09-25"}
[user] tool_result toolu_1 "[hr-001] Paid time off: ... Unused PTO up to 5 days rolls over. ..."
tool_result toolu_2 {"employee": "alice", "as_of": "2026-09-25", "days_left": 12}
[assistant] Alice has 12 PTO days left. Up to 5 unused days roll over to next year [hr-001].
(stop_reason: end_turn)
Two model calls, two tools run in parallel, one answer. ReAct ("reason + act") is the research name for this pattern of alternating a bit of reasoning ("I'll check both at once") with actions (tool calls). The loop itself is ten lines of code. Everything else in this file exists because the ten-line version fails in production.
flowchart LR G[Goal] --> T[Think<br/>choose next step] T --> C[Call a tool] C --> O[Observe result] O --> D{Done, or budget<br/>or step limit hit?} D -->|No| T D -->|Yes| F[Final answer<br/>or hand to human]
Reading it: start at Goal on the left and follow the arrows around the cycle. The model only ever does the Think box: it looks at everything so far and picks the next action. Call a tool and Observe result are your code. The diamond is the only exit, and it has three ways out: the model says it's done, a limit trips, or the agent hands the task to a person. An agent without that diamond is a loop you can't stop.
In code: run_agent is this loop, and it returns an AgentResult
holding the final text, the stop cause, every step, the usage and the cost.
search_kb and get_pto_balance are the two tools in the example.
One round trip, message by message
sequenceDiagram participant U as User participant L as Your loop participant M as Model participant T as Tools U->>L: task L->>M: system + messages + tool definitions M-->>L: tool_use blocks (stop_reason = tool_use) L->>L: validate, check budget, check for loops L->>T: run the calls (concurrently) T-->>L: results or errors L->>M: ONE user message with every tool_result M-->>L: final text (stop_reason = end_turn) L-->>U: answer + stop cause
Reading it: time runs top to bottom. The model's reply is a request
(tool_use), never an action. Everything between that request and the next
call to the model happens in your loop, which is where every control in
this file lives: validation, budgets, loop detection, concurrency. Notice
the model is called twice for one question. Each call re-sends the whole
history, and that's where the cost comes from.
In code: each pass through run_agent appends the model's full
assistant content, runs the requested tools, and appends every result in one
user message; a Step records that one model call and the ToolExecutions
it triggered.
The controls, in the order the loop checks them
Everyday picture. Lending your car to a new driver: a full tank but no more fuel money (budget), "come back by six" (step limit), "if you drive past the same petrol station three times, you're lost, so come home" (loop detection), and "call me if anything feels off" (hand-off to a human).
flowchart TD R[Model response] --> S1{stop_reason?} S1 -->|refusal| X1[stop: refusal] S1 -->|max_tokens| X2[stop: max_tokens<br/>don't run half-written calls] S1 -->|end_turn| X3[stop: completed] S1 -->|tool_use| B{Over token or<br/>dollar budget?} B -->|yes| X4[stop: budget_exceeded] B -->|no| H{Called handoff_to_human?} H -->|yes| X5[stop: handoff] H -->|no| L{Same tool + same args<br/>seen N times?} L -->|yes| X6[stop: loop_detected] L -->|no| E[Run tools concurrently] E --> A[Append all results<br/>in one user message] A --> N{Step limit hit?} N -->|yes| X7[stop: max_steps] N -->|no| M[Call the model again]
Reading it: every box that starts with stop: is an exit, and each of
the seven is a named stop_cause in AgentResult.
Read top to bottom: first trust the model's own stop reason, then protect
your wallet (budget), then honor an explicit escalation, then protect
against repetition, and only then spend time running tools. The order
matters. A truncated response (max_tokens) is checked before tools run,
because a call cut off mid-argument must never execute.
| Failure | What happens | Control in run_agent |
|---|---|---|
| Model never says "done" | Runs forever | max_steps |
| Each turn re-sends the whole conversation | Cost grows quadratically with steps | max_total_tokens, max_cost_usd |
| Model repeats the same call | Burns money, no progress | loop detection on (tool, canonical args) |
| Tool raises | Loop crashes, or the model never learns why | errors become is_error tool results with actionable text |
| Unknown tool / hallucinated name | KeyError |
is_error result listing the real tools |
| Several independent calls in one turn | Slow if run one by one | run concurrently, return all results in one user message |
Output cut off (max_tokens) |
Half-written tool call | stop with cause max_tokens |
Safety refusal (refusal) |
Content is not an answer | stop with cause refusal |
| Task is outside the agent's authority | Agent guesses | a handoff_to_human tool that ends the run cleanly |
"It stopped" is not an outcome you can monitor; "it stopped because of loop_detected at step 3" is.
In code: AgentConfig holds every limit (steps, tokens, dollars, the
loop threshold) and the prices AgentConfig.cost uses to turn primer.agents.llm.Usage into
dollars. A tool raises ToolError to send the model an actionable error
result, and HANDOFF_TOOL_DEF is the hand-off tool's definition.
Parallel tool calls: fan out, fan in
Everyday picture. Three errands in three different shops. You can do them one after another, or send three friends at once and be done when the slowest one gets back.
flowchart LR M[Assistant turn<br/>tool_use A, tool_use B, tool_use C] --> A[run A] M --> B[run B] M --> C[run C] A --> J[One user message:<br/>tool_result A, B, C<br/>matched by tool_use_id] B --> J C --> J
Reading it: the model asked for A, B and C in the same turn, before seeing any result, so they can't depend on each other and it's safe to run them at once. Wall-clock time becomes the slowest call instead of the sum. On the way back, all three results go in a single user message. The API pairs each result with its request by id, and splitting them teaches the model to stop asking for parallel calls.
In code: execute_tools runs one turn's calls on a thread pool and
returns their results in the order they were asked for, turning an unknown
tool name or a raised exception into an error result instead of a crash.
Why cost grows quadratically
Each call sends the entire conversation so far. If every step adds about $t$ tokens, call $k$ sends roughly $k\,t$ input tokens, so a run of $n$ steps sends
Level 3: the formula and its symbols
$$ \text{total input} \approx \sum_{k=1}^{n} k\,t = \frac{t\,n(n+1)}{2} \approx \frac{t\,n^2}{2}. $$
Symbols
| Symbol | Meaning |
|---|---|
| $t$ | new tokens each step adds to the conversation (tool call + result) |
| $k$ | the step number, 1, 2, 3, ... |
| $n$ | total steps in the run |
In words: step $k$ re-sends everything from the $k$ steps before it, so the total is $t$ times $1 + 2 + \dots + n$, which is about half of $n$ squared times $t$.
On the example: with $t = 500$ and $n = 10$, the total is $500 \times 55 = 27{,}500$ input tokens, not the $5{,}000$ you'd guess. Double to $n = 20$ and it's $500 \times 210 = 105{,}000$, almost 4x.
Level 3: in Python
In Python:
t = 500
def total_input(n):
# Σ_k k·t: step k re-sends k steps' worth
return sum(k * t for k in range(1, n + 1))
total_input(10) # → 27500
total_input(20) # → 105000
# almost 4x
round(total_input(20) / total_input(10), 1) # → 3.8
Reading it: each bar is one call to the model in the budget-burner demo (a model that keeps searching). Bars grow step by step because the history grows. The line is the running total, which is what you pay for, and it bends upward. The dashed horizontal line is the token budget. The run stops at the first step where the total crosses it, instead of carrying on.
Reading it: the x-axis is the number of steps in a run, and the y-axis is
the total input tokens sent. The straight line is what people intuitively expect
("each step costs the same"). The curve is what actually happens when every
call re-sends the history. At 10 steps the gap is about 5x, and it keeps
widening. That gap is why step limits, trimming tool output, and prompt
caching (primer.agents.cost) matter.
Reading it: the x-axis is how many tools the model requested in one turn, each with a realistic latency between 0.2 s and 1.2 s. Run one after another, the time is the sum and climbs with every tool. Run concurrently, it's the max, which flattens out near the slowest single call. Parallel tool calls are among the cheapest latency wins in agent systems.
In code: cumulative_input_tokens evaluates the formula above, $t$
times $n(n+1)/2$, for any run length.
Running it against a real model
The loop only depends on the LLM protocol, so the same code runs against
Claude:
from primer.agents.llm import ClaudeLLM
from primer.agents.agent_loop import run_agent, AgentConfig
result = run_agent(
ClaudeLLM(), # needs `pip install anthropic` + credentials
task="How many PTO days does alice have left, and what's the rollover rule?",
tools=HR_TOOLS, # name -> python function
tool_defs=HR_TOOL_DEFS, # Anthropic tool definitions (JSON Schema)
system="You are an HR assistant. Use tools; cite doc ids.",
config=AgentConfig(max_steps=8, max_cost_usd=0.50),
)
print(result.stop_cause, result.final_text)
The SDK also ships a tool runner (client.beta.messages.tool_runner)
that drives this loop for you. Owning the loop, as here, is what you do
when you need custom budgets, loop detection, approval gates or tracing in
exactly your shape.
In 20 seconds
- The agent loop is: call model, run the tools it asks for, append results, repeat until
end_turn. - The model never executes anything. Your code validates and runs every tool call.
- Production loops need step limits, token/cost budgets, loop detection, error-as-result handling, and a human hand-off.
- Parallel tool calls: run them concurrently, return all results in a single user message.
- Input cost grows roughly with the square of the number of steps, because every call re-sends the history.
Self-test questions
Q: An agent in production suddenly costs 5x more per task. What do you check first? A: Steps per task and tokens per step, from traces. A jump usually means a loop (the same call repeated), a tool that started returning huge payloads, or a prompt change that stopped the model from recognizing "done". Step budgets and loop detection cap the damage; alerts on tokens-per-task catch it early.
Q: A tool throws an exception mid-run. What should the loop do?
A: Catch it and return a tool_result with is_error: true and a message the
model can act on ("date must be YYYY-MM-DD, got 'next friday'"). The model
usually fixes its next call. Crashing the loop throws away the work so far;
swallowing the error silently makes the model guess.
Q: Why return all parallel tool results in one message?
A: The API pairs each tool_use with its tool_result by id in the next user
turn. Splitting results across several messages breaks that pairing and
teaches the model to stop making parallel calls, which slows every run.
Q: When should an agent hand off to a human? A: When it lacks authority (refunds over a limit), lacks information after reasonable search, detects conflicting sources, or hits a budget. Make hand-off a tool, so it is an explicit, logged outcome rather than a vague final answer.
The papers behind this lesson
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022). https://arxiv.org/abs/2210.03629. It showed that interleaving short reasoning traces with tool actions, then observing the results, beats reasoning alone or acting alone, and this loop is the shape of nearly every agent built since. annotated companion
Further reading
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
- Claude tool use, implementing the loop: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022): https://arxiv.org/abs/2210.03629
- Lilian Weng, LLM Powered Autonomous Agents: https://lilianweng.github.io/posts/2023-06-23-agent/
1r""" 2# The agent loop, production-grade 3 4Run: `python -m primer.agents.agent_loop` 5 6New to the notation? `primer.notation` explains every symbol used here from 7zero. This lesson builds on the message format and tool calling in 8`primer.agents.llm` and on the autonomy spectrum in 9`primer.agents.orchestration`. 10 11## Level 1: The practitioner's guide 12 13**In one sentence.** The agent loop is the piece of code that calls the 14model, runs the tools it asks for, sends the results back and repeats until 15the model says it is done or a limit you set says it is done for it. 16 17**When you need it.** You need a loop whenever the model has to take more 18than one step whose order it decides: look something up, then act on what 19it found, then check. One call with one tool call you run once is not a 20loop, and neither is a fixed pipeline where your code decides the steps 21(`primer.agents.orchestration`). You need the *production* loop, rather 22than the ten-line version, the moment real money and real side effects are 23attached, because the ten-line version has no way to stop. In this 24lesson's happy path, one question ("How many PTO days does alice have 25left, and what rolls over?") takes 2 model calls and 2 tool calls run in 26parallel, and stops because the model said `end_turn`. In the same 27lesson, a model that keeps searching stops only because a token budget 28tripped: 11 calls in, the running total passes 12,000 tokens (12,990) and 29the run ends with an estimated cost of \$0.07. The tell that you need the 30controls: you cannot say, for every run that ended, *why* it ended. 31 32**Your options.** Six ways to run the loop, from the least machinery to the 33most: 34 35| Option | What it does | What it guarantees | What it costs | Where it lives | 36|---|---|---|---|---| 37| No loop | One model call; if it asks for a tool, run it once and call once more | Bounded cost: at most two calls | Nothing that needs a second decision gets done | Your code | 38| The SDK's tool runner | Drives the request, run, reply cycle until the model returns no tool call or `max_iterations` is reached | The plumbing done right: results paired by id, one message per turn, type-checked inputs | Per-turn controls are limited; Claude's docs point you to the manual loop for approval, custom logging or conditional execution | The provider's SDK (seven languages for Claude) | 39| The ten-line loop | Call, run tools, append results, repeat while `stop_reason` is `tool_use` | Full control of every turn | Every production failure is yours: it can run forever, repeat itself, crash on a tool error or execute a half-written call | Your code | 40| The production loop (this lesson) | The ten-line loop plus a step limit, token and dollar budgets, loop detection, errors returned as results, a hand-off tool and parallel tool execution | A named stop cause for every run, one of seven | A config to tune (steps, tokens, dollars, repeat threshold) and a test per control | Your code | 41| A framework's loop | The loop with tracing, persistence and hand-offs built in, and its own limit (`max_turns`, a recursion limit) | Standard patterns and observability without writing them | A dependency, its assumptions and its defaults (LangGraph's recursion limit defaults to 1,000 steps) | LangGraph, the OpenAI Agents SDK, the Claude Agent SDK | 42| A hosted loop | The provider runs the loop and a sandbox for the tools; you send messages and tool results | No loop code, no state files, per-session containers | Session runtime on top of tokens (Claude Managed Agents lists \$0.08 per session-hour) and less say over each turn | The provider's servers | 43 44**How to choose.** Start from what can go wrong and who pays for it. 45 46- A prototype, a notebook, a one-off script: the tool runner. It gets the 47 message pairing right, which is the part people get wrong first. 48- Production with side effects (writes, emails, refunds): own the loop, or 49 use a framework whose per-turn hooks you have read. You need an approval 50 gate, a budget in dollars, and a hand-off to a person, in exactly your 51 shape. 52- Many independent tool calls per turn (three lookups): make sure whatever 53 you use runs them concurrently and returns all results in one message. 54 The lesson's latency figure shows serial time climbing with every tool 55 while concurrent time flattens at the slowest call. 56- Long-running or scheduled agents you would rather not host: a hosted loop, 57 if its controls cover your approval and budget needs. 58- Whatever you pick, set every limit before the first real run: steps, 59 tokens, dollars and a repeat threshold. This lesson's defaults are 10 60 steps, 50,000 tokens, and a stop when the same tool is asked for with the 61 same arguments 3 times. 62 63**What it costs.** Input tokens grow with the square of the number of 64steps, because each call re-sends the whole history: with 500 new tokens a 65step, 10 steps send 27,500 input tokens rather than the 5,000 a flat 66per-step cost suggests, and 20 steps send 105,000, almost four times ten 67steps' total. At Claude Opus 5's list prices (\$5 per million input tokens 68and \$25 per million output, the defaults in this lesson's `AgentConfig`), 69the runaway run above cost about 7 cents before its 12,000-token budget 70stopped it. Latency is one round trip per model call plus the slowest tool 71in each turn; a serial loop adds every tool's time instead. The controls 72themselves cost nothing per call: loop detection is a dictionary of (tool, 73arguments) counts, and a budget is a comparison. 74 75**What breaks.** 76 77- **The model never says "done".** Without a step limit the loop runs 78 until the money does. Set `max_steps`, and treat hitting it as an outcome 79 to monitor, not an error to hide. 80- **The same call, again and again.** A model retrying one search burns a 81 call per repeat and never progresses. Count (tool, canonical arguments) 82 and stop at a threshold. 83- **A half-written tool call.** A reply cut off by `max_tokens` can end 84 mid-argument. Check the stop reason before running anything; a truncated 85 call must never execute. 86- **A tool throws.** Crashing the loop discards the work so far; swallowing 87 the error makes the model guess. Return the failure as an error result 88 with a message the model can act on ("as_of must be YYYY-MM-DD"); in the 89 lesson, the model fixes its own call on the next step. 90- **A tool the model made up.** An unknown name is a `KeyError` in a naive 91 loop. Answer with an error result listing the real tools. 92- **Results split across messages.** The API pairs each request and result 93 by id inside one user turn; splitting them breaks the pairing and teaches 94 the model to stop making parallel calls. 95- **Guessing outside its authority.** A refund over the limit needs a 96 person. Make hand-off a tool, so it ends the run with a logged reason 97 instead of a confident wrong answer. 98 99**In the wild.** ReAct (Yao et al., 2022) named the pattern the loop 100implements: a short piece of reasoning, an action, an observation, repeated. 101Claude's SDKs ship a tool runner that loops until the model returns no tool 102use or `max_iterations` is reached, and their docs send you to the manual 103loop when you need approval or custom logging. The OpenAI Agents SDK runs 104the same call, classify, run-tools cycle and raises `MaxTurnsExceeded` past 105`max_turns`. LangGraph bounds a graph with a recursion limit and raises 106`GraphRecursionError` when it trips. Anthropic's *Building effective agents* 107adds the operational advice: agents trade cost and latency for open-ended 108capability, so test them in sandboxes and put guardrails in. Every one of 109these products is this lesson's diamond ("done, or budget, or step limit?") 110with a different name on it. 111 112**Go deeper.** Level 2 traces one round trip message by message, orders the 113seven exits the way the loop checks them and says why that order matters, 114fans tool calls out and back in, and derives the quadratic cost with a 115formula you can rerun against the budget-burner demo. If you only needed to 116choose, you are done. 117 118## Level 2: How it works, from scratch 119 120The loop itself is ten lines: call the model, run the tools it asks for, 121append the results, repeat until it stops asking. Everything else in this 122lesson exists because the ten-line version fails in production, and each 123failure gets a named control. 124 125## The idea 126 127**Everyday picture.** A new assistant runs errands for you. They can't do 128anything themselves. They come back after *every* errand and say: "here's 129what I found; next I'd like to do X." You carry out X and hand them the 130result. They decide the next step, and so on, until they say "done" (or you 131say "that's enough, you've spent the budget"). The assistant is the model, 132the errands are tool calls, and *you* are the loop in this file. 133 134**Tiny worked example.** "How many PTO days does alice have left, and what 135rolls over?" Here's the real transcript the happy-path demo produces: 136 137```text 138[user] How many PTO days does alice have left, and what rolls over? 139[assistant] I'll check the PTO policy and alice's balance at the same time. 140 tool_use toolu_1 search_kb {"query": "PTO rollover"} 141 tool_use toolu_2 get_pto_balance {"employee": "alice", "as_of": "2026-09-25"} 142[user] tool_result toolu_1 "[hr-001] Paid time off: ... Unused PTO up to 5 days rolls over. ..." 143 tool_result toolu_2 {"employee": "alice", "as_of": "2026-09-25", "days_left": 12} 144[assistant] Alice has 12 PTO days left. Up to 5 unused days roll over to next year [hr-001]. 145 (stop_reason: end_turn) 146``` 147 148Two model calls, two tools run in parallel, one answer. **ReAct** ("reason + 149act") is the research name for this pattern of alternating a bit of 150reasoning ("I'll check both at once") with actions (tool calls). The loop 151itself is ten lines of code. Everything else in this file exists because 152the ten-line version fails in production. 153 154```mermaid 155flowchart LR 156 G[Goal] --> T[Think<br/>choose next step] 157 T --> C[Call a tool] 158 C --> O[Observe result] 159 O --> D{Done, or budget<br/>or step limit hit?} 160 D -->|No| T 161 D -->|Yes| F[Final answer<br/>or hand to human] 162``` 163 164**Reading it:** start at *Goal* on the left and follow the arrows around the 165cycle. The model only ever does the *Think* box: it looks at everything so 166far and picks the next action. *Call a tool* and *Observe result* are your 167code. The diamond is the only exit, and it has three ways out: the model 168says it's done, a limit trips, or the agent hands the task to a person. An 169agent without that diamond is a loop you can't stop. 170 171**In code:** `run_agent` is this loop, and it returns an `AgentResult` 172holding the final text, the stop cause, every step, the usage and the cost. 173`search_kb` and `get_pto_balance` are the two tools in the example. 174 175## One round trip, message by message 176 177```mermaid 178sequenceDiagram 179 participant U as User 180 participant L as Your loop 181 participant M as Model 182 participant T as Tools 183 U->>L: task 184 L->>M: system + messages + tool definitions 185 M-->>L: tool_use blocks (stop_reason = tool_use) 186 L->>L: validate, check budget, check for loops 187 L->>T: run the calls (concurrently) 188 T-->>L: results or errors 189 L->>M: ONE user message with every tool_result 190 M-->>L: final text (stop_reason = end_turn) 191 L-->>U: answer + stop cause 192``` 193 194**Reading it:** time runs top to bottom. The model's reply is a *request* 195(`tool_use`), never an action. Everything between that request and the next 196call to the model happens in your loop, which is where every control in 197this file lives: validation, budgets, loop detection, concurrency. Notice 198the model is called twice for one question. Each call re-sends the whole 199history, and that's where the cost comes from. 200 201**In code:** each pass through `run_agent` appends the model's full 202assistant content, runs the requested tools, and appends every result in one 203user message; a `Step` records that one model call and the `ToolExecution`s 204it triggered. 205 206## The controls, in the order the loop checks them 207 208**Everyday picture.** Lending your car to a new driver: a full tank but no 209more fuel money (budget), "come back by six" (step limit), "if you drive past 210the same petrol station three times, you're lost, so come home" (loop 211detection), and "call me if anything feels off" (hand-off to a human). 212 213```mermaid 214flowchart TD 215 R[Model response] --> S1{stop_reason?} 216 S1 -->|refusal| X1[stop: refusal] 217 S1 -->|max_tokens| X2[stop: max_tokens<br/>don't run half-written calls] 218 S1 -->|end_turn| X3[stop: completed] 219 S1 -->|tool_use| B{Over token or<br/>dollar budget?} 220 B -->|yes| X4[stop: budget_exceeded] 221 B -->|no| H{Called handoff_to_human?} 222 H -->|yes| X5[stop: handoff] 223 H -->|no| L{Same tool + same args<br/>seen N times?} 224 L -->|yes| X6[stop: loop_detected] 225 L -->|no| E[Run tools concurrently] 226 E --> A[Append all results<br/>in one user message] 227 A --> N{Step limit hit?} 228 N -->|yes| X7[stop: max_steps] 229 N -->|no| M[Call the model again] 230``` 231 232**Reading it:** every box that starts with `stop:` is an exit, and each of 233the seven is a named `stop_cause` in `AgentResult`. 234Read top to bottom: first trust the model's own stop reason, then protect 235your wallet (budget), then honor an explicit escalation, then protect 236against repetition, and only then spend time running tools. The order 237matters. A truncated response (`max_tokens`) is checked before tools run, 238because a call cut off mid-argument must never execute. 239 240| Failure | What happens | Control in `run_agent` | 241|---|---|---| 242| Model never says "done" | Runs forever | `max_steps` | 243| Each turn re-sends the whole conversation | Cost grows quadratically with steps | `max_total_tokens`, `max_cost_usd` | 244| Model repeats the same call | Burns money, no progress | loop detection on (tool, canonical args) | 245| Tool raises | Loop crashes, or the model never learns why | errors become `is_error` tool results with actionable text | 246| Unknown tool / hallucinated name | `KeyError` | `is_error` result listing the real tools | 247| Several independent calls in one turn | Slow if run one by one | run concurrently, return **all** results in **one** user message | 248| Output cut off (`max_tokens`) | Half-written tool call | stop with cause `max_tokens` | 249| Safety refusal (`refusal`) | Content is not an answer | stop with cause `refusal` | 250| Task is outside the agent's authority | Agent guesses | a `handoff_to_human` tool that ends the run cleanly | 251 252"It stopped" is not an outcome you can monitor; "it stopped because of 253loop_detected at step 3" is. 254 255**In code:** `AgentConfig` holds every limit (steps, tokens, dollars, the 256loop threshold) and the prices `AgentConfig.cost` uses to turn `primer.agents.llm.Usage` into 257dollars. A tool raises `ToolError` to send the model an actionable error 258result, and `HANDOFF_TOOL_DEF` is the hand-off tool's definition. 259 260## Parallel tool calls: fan out, fan in 261 262**Everyday picture.** Three errands in three different shops. You can do them 263one after another, or send three friends at once and be done when the 264slowest one gets back. 265 266```mermaid 267flowchart LR 268 M[Assistant turn<br/>tool_use A, tool_use B, tool_use C] --> A[run A] 269 M --> B[run B] 270 M --> C[run C] 271 A --> J[One user message:<br/>tool_result A, B, C<br/>matched by tool_use_id] 272 B --> J 273 C --> J 274``` 275 276**Reading it:** the model asked for A, B and C in the same turn, before 277seeing any result, so they can't depend on each other and it's safe to run 278them at once. Wall-clock time becomes the *slowest* call instead of the 279*sum*. On the way back, all three results go in a single user message. The 280API pairs each result with its request by id, and splitting them teaches the 281model to stop asking for parallel calls. 282 283**In code:** `execute_tools` runs one turn's calls on a thread pool and 284returns their results in the order they were asked for, turning an unknown 285tool name or a raised exception into an error result instead of a crash. 286 287## Why cost grows quadratically 288 289Each call sends the *entire* conversation so far. If every step adds about 290$t$ tokens, call $k$ sends roughly $k\,t$ input tokens, so a run of $n$ steps 291sends 292 293$$ 294\text{total input} \approx \sum_{k=1}^{n} k\,t = \frac{t\,n(n+1)}{2} \approx \frac{t\,n^2}{2}. 295$$ 296 297**Symbols** 298 299| Symbol | Meaning | 300|---|---| 301| $t$ | new tokens each step adds to the conversation (tool call + result) | 302| $k$ | the step number, 1, 2, 3, ... | 303| $n$ | total steps in the run | 304 305**In words:** step $k$ re-sends everything from the $k$ steps before it, so 306the total is $t$ times $1 + 2 + \dots + n$, which is about half of $n$ squared times $t$. 307 308**On the example:** with $t = 500$ and $n = 10$, the total is 309$500 \times 55 = 27{,}500$ input tokens, not the $5{,}000$ you'd guess. Double 310to $n = 20$ and it's $500 \times 210 = 105{,}000$, almost 4x. 311 312**In Python:** 313 314```python 315t = 500 316def total_input(n): 317 # Σ_k k·t: step k re-sends k steps' worth 318 return sum(k * t for k in range(1, n + 1)) 319total_input(10) # → 27500 320total_input(20) # → 105000 321# almost 4x 322round(total_input(20) / total_input(10), 1) # → 3.8 323``` 324 325 326 327**Reading it:** each bar is one call to the model in the budget-burner demo 328(a model that keeps searching). Bars grow step by step because the history 329grows. The line is the running total, which is what you pay for, and it 330bends upward. The dashed horizontal line is the token budget. The run stops 331at the first step where the total crosses it, instead of carrying on. 332 333 334 335**Reading it:** the x-axis is the number of steps in a run, and the y-axis is 336the total input tokens sent. The straight line is what people intuitively expect 337("each step costs the same"). The curve is what actually happens when every 338call re-sends the history. At 10 steps the gap is about 5x, and it keeps 339widening. That gap is why step limits, trimming tool output, and prompt 340caching (`primer.agents.cost`) matter. 341 342 343 344**Reading it:** the x-axis is how many tools the model requested in one turn, 345each with a realistic latency between 0.2 s and 1.2 s. Run one after another, 346the time is the *sum* and climbs with every tool. Run concurrently, it's 347the *max*, which flattens out near the slowest single call. Parallel tool 348calls are among the cheapest latency wins in agent systems. 349 350**In code:** `cumulative_input_tokens` evaluates the formula above, $t$ 351times $n(n+1)/2$, for any run length. 352 353## Running it against a real model 354 355The loop only depends on the `LLM` protocol, so the same code runs against 356Claude: 357 358```python 359from primer.agents.llm import ClaudeLLM 360from primer.agents.agent_loop import run_agent, AgentConfig 361 362result = run_agent( 363 ClaudeLLM(), # needs `pip install anthropic` + credentials 364 task="How many PTO days does alice have left, and what's the rollover rule?", 365 tools=HR_TOOLS, # name -> python function 366 tool_defs=HR_TOOL_DEFS, # Anthropic tool definitions (JSON Schema) 367 system="You are an HR assistant. Use tools; cite doc ids.", 368 config=AgentConfig(max_steps=8, max_cost_usd=0.50), 369) 370print(result.stop_cause, result.final_text) 371``` 372 373The SDK also ships a *tool runner* (`client.beta.messages.tool_runner`) 374that drives this loop for you. Owning the loop, as here, is what you do 375when you need custom budgets, loop detection, approval gates or tracing in 376exactly your shape. 377 378## In 20 seconds 379- The agent loop is: call model, run the tools it asks for, append results, repeat until `end_turn`. 380- The model never executes anything. Your code validates and runs every tool call. 381- Production loops need step limits, token/cost budgets, loop detection, error-as-result handling, and a human hand-off. 382- Parallel tool calls: run them concurrently, return all results in a single user message. 383- Input cost grows roughly with the square of the number of steps, because every call re-sends the history. 384 385## Self-test questions 386 387**Q: An agent in production suddenly costs 5x more per task. What do you check first?** 388A: Steps per task and tokens per step, from traces. A jump usually means a loop 389(the same call repeated), a tool that started returning huge payloads, or a 390prompt change that stopped the model from recognizing "done". Step budgets and 391loop detection cap the damage; alerts on tokens-per-task catch it early. 392 393**Q: A tool throws an exception mid-run. What should the loop do?** 394A: Catch it and return a `tool_result` with `is_error: true` and a message the 395model can act on ("date must be YYYY-MM-DD, got 'next friday'"). The model 396usually fixes its next call. Crashing the loop throws away the work so far; 397swallowing the error silently makes the model guess. 398 399**Q: Why return all parallel tool results in one message?** 400A: The API pairs each `tool_use` with its `tool_result` by id in the next user 401turn. Splitting results across several messages breaks that pairing and 402teaches the model to stop making parallel calls, which slows every run. 403 404**Q: When should an agent hand off to a human?** 405A: When it lacks authority (refunds over a limit), lacks information after 406reasonable search, detects conflicting sources, or hits a budget. Make hand-off a 407tool, so it is an explicit, logged outcome rather than a vague final answer. 408 409## The papers behind this lesson 410 411- **Yao et al., *ReAct: Synergizing Reasoning and Acting in Language Models* (2022).** 412 https://arxiv.org/abs/2210.03629. It showed that interleaving short reasoning 413 traces with tool actions, then observing the results, beats reasoning alone or 414 acting alone, and this loop is the shape of nearly every agent built since. 415 [annotated companion](../../papers/react.html) 416 417## Further reading 418- Anthropic, *Building effective agents*: https://www.anthropic.com/engineering/building-effective-agents 419- Claude tool use, implementing the loop: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use 420- Yao et al., *ReAct: Synergizing Reasoning and Acting in Language Models* (2022): https://arxiv.org/abs/2210.03629 421- Lilian Weng, *LLM Powered Autonomous Agents*: https://lilianweng.github.io/posts/2023-06-23-agent/ 422""" 423 424from __future__ import annotations 425 426import json 427from collections import Counter 428from concurrent.futures import ThreadPoolExecutor 429from dataclasses import dataclass, field 430from typing import Any, Callable, Literal 431 432from primer._show import banner, say, table, takeaway 433from primer.agents.llm import ( 434 LLM, 435 LLMResponse, 436 ScriptedLLM, 437 ToolCall, 438 Usage, 439 last_user_text, 440 tool_result_block, 441 tool_results, 442) 443 444StopCause = Literal[ 445 "completed", # model said end_turn with a final answer 446 "max_steps", # step limit hit 447 "budget_exceeded", # token or dollar budget hit 448 "loop_detected", # same tool + same args repeated too often 449 "max_tokens", # a response was truncated 450 "refusal", # the model declined 451 "handoff", # the model handed the task to a human 452] 453 454HANDOFF_TOOL = "handoff_to_human" 455 456# The hand-off tool definition. Add it to `tool_defs` to give the agent an 457# explicit, auditable way to say "this needs a person". 458HANDOFF_TOOL_DEF: dict[str, Any] = { 459 "name": HANDOFF_TOOL, 460 "description": ( 461 "Escalate to a human when the request is outside your authority (for example refunds, " 462 "policy exceptions, legal questions), when sources conflict, or when you cannot find the " 463 "answer after searching. Do not use for questions you can answer with the other tools." 464 ), 465 "input_schema": { 466 "type": "object", 467 "properties": {"reason": {"type": "string", "description": "Why a human is needed, one sentence."}}, 468 "required": ["reason"], 469 }, 470} 471 472 473class ToolError(Exception): 474 """Raise from a tool to send an actionable `is_error` result back to the model. 475 476 The message is shown to the model verbatim, so write it for the model: 477 say what was wrong and what a valid call looks like. 478 """ 479 480 481@dataclass 482class AgentConfig: 483 max_steps: int = 10 484 max_total_tokens: int = 50_000 485 max_cost_usd: float | None = None 486 # Stop when the same (tool, args) pair has been requested this many times. 487 loop_threshold: int = 3 488 max_parallel_tools: int = 4 489 max_tokens_per_call: int = 4096 490 # Dollar prices per million tokens. Defaults are Claude Opus 5 list prices 491 # at the time of writing; check the provider's pricing page. 492 price_in_per_mtok: float = 5.0 493 price_out_per_mtok: float = 25.0 494 495 def cost(self, usage: Usage) -> float: 496 return (usage.input_tokens * self.price_in_per_mtok + usage.output_tokens * self.price_out_per_mtok) / 1e6 497 498 499@dataclass 500class ToolExecution: 501 name: str 502 input: dict[str, Any] 503 content: str 504 is_error: bool 505 506 507@dataclass 508class Step: 509 """One model call plus the tools it triggered.""" 510 511 index: int 512 text: str 513 tool_executions: list[ToolExecution] 514 usage: Usage 515 stop_reason: str 516 517 518@dataclass 519class AgentResult: 520 final_text: str 521 stop_cause: StopCause 522 steps: list[Step] 523 usage: Usage 524 cost_usd: float 525 transcript: list[dict[str, Any]] = field(repr=False) 526 detail: str = "" # human-readable explanation of the stop cause 527 528 @property 529 def ok(self) -> bool: 530 return self.stop_cause == "completed" 531 532 @property 533 def tool_call_count(self) -> int: 534 return sum(len(s.tool_executions) for s in self.steps) 535 536 537def _signature(call: ToolCall) -> str: 538 # Canonical JSON (sorted keys) so {"a":1,"b":2} and {"b":2,"a":1} count as the same call. 539 return f"{call.name}:{json.dumps(call.input, sort_keys=True, default=str)}" 540 541 542def _execute_one(call: ToolCall, tools: dict[str, Callable[..., Any]]) -> ToolExecution: 543 """Run one tool call, turning every failure into an actionable error result.""" 544 fn = tools.get(call.name) 545 if fn is None: 546 # Hallucinated tool names happen. Tell the model what exists. 547 return ToolExecution(call.name, call.input, f"Unknown tool '{call.name}'. Available tools: {sorted(tools)}.", True) 548 try: 549 out = fn(**call.input) 550 return ToolExecution(call.name, call.input, out if isinstance(out, str) else json.dumps(out, default=str), False) 551 except ToolError as e: 552 return ToolExecution(call.name, call.input, str(e), True) 553 except TypeError as e: 554 # Usually wrong/missing argument names. The message names the bad argument. 555 return ToolExecution(call.name, call.input, f"Bad arguments for {call.name}: {e}", True) 556 except Exception as e: # noqa: BLE001 557 # Don't leak stack traces or internals into the context; give the class and message. 558 return ToolExecution(call.name, call.input, f"{call.name} failed ({type(e).__name__}): {e}", True) 559 560 561def execute_tools(calls: list[ToolCall], tools: dict[str, Callable[..., Any]], max_workers: int = 4) -> list[ToolExecution]: 562 """Run tool calls concurrently. Results come back in the same order as `calls`. 563 564 Tools in one assistant turn are independent by construction (the model 565 asked for them together, before seeing any result), so running them in 566 parallel is safe and cuts wall-clock time to the slowest call. 567 """ 568 if len(calls) <= 1: 569 return [_execute_one(c, tools) for c in calls] 570 with ThreadPoolExecutor(max_workers=max_workers) as pool: 571 return list(pool.map(lambda c: _execute_one(c, tools), calls)) 572 573 574def run_agent( 575 llm: LLM, 576 task: str, 577 tools: dict[str, Callable[..., Any]], 578 tool_defs: list[dict[str, Any]], 579 system: str = "You are a helpful assistant. Use the tools when needed.", 580 config: AgentConfig | None = None, 581) -> AgentResult: 582 """Run the think -> act -> observe loop until done or a control trips. 583 584 Args: 585 llm: anything implementing `LLM.complete` (ScriptedLLM or ClaudeLLM). 586 task: the user's request. 587 tools: tool name -> Python callable taking the tool input as kwargs. 588 tool_defs: Anthropic-format tool definitions sent to the model. 589 system: system prompt. 590 config: limits and prices. 591 592 Returns: 593 AgentResult with the final text, stop cause, per-step record, usage, 594 cost and the full transcript (for tracing and replay). 595 """ 596 config = config or AgentConfig() 597 messages: list[dict[str, Any]] = [{"role": "user", "content": task}] 598 steps: list[Step] = [] 599 usage = Usage() 600 seen: Counter[str] = Counter() 601 602 def finish(cause: StopCause, text: str = "", detail: str = "") -> AgentResult: 603 return AgentResult(text, cause, steps, usage, config.cost(usage), messages, detail) 604 605 for i in range(config.max_steps): 606 resp: LLMResponse = llm.complete( 607 system=system, messages=messages, tools=tool_defs, max_tokens=config.max_tokens_per_call 608 ) 609 usage += resp.usage 610 # Always append the full assistant content (it may carry thinking blocks 611 # that must be sent back unchanged with the real API). 612 messages.append({"role": "assistant", "content": resp.assistant_content}) 613 step = Step(i, resp.text, [], resp.usage, resp.stop_reason) 614 steps.append(step) 615 616 # --- Stop reasons that aren't "use a tool" ------------------------- 617 if resp.stop_reason == "refusal": 618 return finish("refusal", resp.text, "model declined the request") 619 if resp.stop_reason == "max_tokens": 620 # A truncated response may contain a half-formed tool call. Don't run it. 621 return finish("max_tokens", resp.text, "response truncated; raise max_tokens or ask for less output") 622 if resp.stop_reason == "end_turn" or not resp.tool_calls: 623 return finish("completed", resp.text) 624 625 # --- Budget: checked after every call, before spending more --------- 626 total = usage.input_tokens + usage.output_tokens 627 if total > config.max_total_tokens: 628 return finish("budget_exceeded", resp.text, f"{total:,} tokens > budget {config.max_total_tokens:,}") 629 if config.max_cost_usd is not None and config.cost(usage) > config.max_cost_usd: 630 return finish("budget_exceeded", resp.text, f"${config.cost(usage):.4f} > budget ${config.max_cost_usd:.4f}") 631 632 # --- Hand-off: an explicit, logged outcome -------------------------- 633 for call in resp.tool_calls: 634 if call.name == HANDOFF_TOOL: 635 return finish("handoff", call.input.get("reason", ""), "escalated to a human") 636 637 # --- Loop detection: same call requested too many times ------------- 638 for call in resp.tool_calls: 639 seen[_signature(call)] += 1 640 if seen[_signature(call)] >= config.loop_threshold: 641 return finish( 642 "loop_detected", resp.text, f"{call.name}({json.dumps(call.input)}) requested {seen[_signature(call)]} times" 643 ) 644 645 # --- Act: run all requested tools concurrently ---------------------- 646 executions = execute_tools(resp.tool_calls, tools, config.max_parallel_tools) 647 step.tool_executions = executions 648 # All results go back in ONE user message, matched by tool_use_id. 649 messages.append( 650 { 651 "role": "user", 652 "content": [ 653 tool_result_block(call.id, ex.content, ex.is_error) for call, ex in zip(resp.tool_calls, executions) 654 ], 655 } 656 ) 657 658 return finish("max_steps", steps[-1].text if steps else "", f"hit max_steps={config.max_steps}") 659 660 661# --------------------------------------------------------------------------- 662# Demo tools and scripted "models" 663# --------------------------------------------------------------------------- 664 665PTO_BALANCES = {"alice": 12, "bob": 3} 666 667SEARCH_KB_DEF = { 668 "name": "search_kb", 669 "description": "Search the company knowledge base (IT, HR, finance policies). Returns doc ids and text.", 670 "input_schema": {"type": "object", "properties": {"query": {"type": "string"}}, "required": ["query"]}, 671} 672PTO_BALANCE_DEF = { 673 "name": "get_pto_balance", 674 "description": "Get an employee's remaining PTO days as of a date. Use for 'how many days do I have left'.", 675 "input_schema": { 676 "type": "object", 677 "properties": { 678 "employee": {"type": "string", "description": "lowercase username, e.g. 'alice'"}, 679 "as_of": {"type": "string", "description": "date as YYYY-MM-DD"}, 680 }, 681 "required": ["employee", "as_of"], 682 }, 683} 684 685 686def search_kb(query: str) -> str: 687 """Tiny keyword search over the shared corpus (see primer.agents.rag for the real thing).""" 688 from primer.common.corpus import DOCS 689 from primer.common.text import tokenize 690 691 q = set(tokenize(query)) 692 scored = sorted(DOCS, key=lambda d: -len(q & set(tokenize(d.title + " " + d.text)))) 693 return "\n".join(f"[{d.id}] {d.title}: {d.text}" for d in scored[:2]) 694 695 696def get_pto_balance(employee: str, as_of: str) -> str: 697 import re 698 699 if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", as_of): 700 # Actionable: says what's wrong AND what a valid value looks like. 701 raise ToolError(f"as_of must be a date in YYYY-MM-DD format (e.g. 2026-10-02), got {as_of!r}.") 702 if employee not in PTO_BALANCES: 703 raise ToolError(f"Unknown employee {employee!r}. Known usernames: {sorted(PTO_BALANCES)}.") 704 return json.dumps({"employee": employee, "as_of": as_of, "days_left": PTO_BALANCES[employee]}) 705 706 707DEMO_TOOLS = {"search_kb": search_kb, "get_pto_balance": get_pto_balance} 708DEMO_TOOL_DEFS = [SEARCH_KB_DEF, PTO_BALANCE_DEF, HANDOFF_TOOL_DEF] 709 710 711def happy_policy(system, messages, tools): 712 """Turn 1: two independent lookups in parallel. Turn 2: answer from the results.""" 713 results = tool_results(messages) 714 if not results: 715 return ( 716 "I'll check the PTO policy and alice's balance at the same time.", 717 [ToolCall("", "search_kb", {"query": "PTO rollover"}), ToolCall("", "get_pto_balance", {"employee": "alice", "as_of": "2026-09-25"})], 718 ) 719 balance = json.loads(results[-1]["content"])["days_left"] 720 return f"Alice has {balance} PTO days left. Up to 5 unused days roll over to next year [hr-001]." 721 722 723def looping_policy(system, messages, tools): 724 """A confused model that keeps re-running the same search.""" 725 return ToolCall("", "search_kb", {"query": "PTO rollover"}) 726 727 728def budget_burner_policy(system, messages, tools): 729 """A model that keeps searching with new queries: no loop, just ever-growing context.""" 730 n = len(tool_results(messages)) 731 return ToolCall("", "search_kb", {"query": f"PTO policy details page {n}"}) 732 733 734def recovering_policy(system, messages, tools): 735 """First call has a bad date; the model reads the actionable error and fixes it.""" 736 results = tool_results(messages) 737 if not results: 738 return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "next friday"}) 739 last = results[-1] 740 if last.get("is_error"): 741 return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "2026-10-02"}) 742 return f"Bob will have {json.loads(last['content'])['days_left']} PTO days left on 2026-10-02." 743 744 745def handoff_policy(system, messages, tools): 746 if "refund" in last_user_text(messages).lower(): 747 return ToolCall("", HANDOFF_TOOL, {"reason": "Refund requests over $500 need a finance approver."}) 748 return "OK." 749 750 751def cumulative_input_tokens(n_steps: int, tokens_per_step: int) -> int: 752 """Total input tokens for an n-step run when step k re-sends k·t tokens of history.""" 753 return tokens_per_step * n_steps * (n_steps + 1) // 2 754 755 756def figures() -> dict: 757 """Plots computed from this lesson's own code. matplotlib is imported here 758 so the lesson itself needs only NumPy.""" 759 import matplotlib 760 761 matplotlib.use("Agg") 762 import matplotlib.pyplot as plt 763 import numpy as np 764 765 figs = {} 766 767 # 1. The budget-burner run, step by step. 768 budget = 12_000 769 run = run_agent( 770 ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS, 771 config=AgentConfig(max_steps=50, max_total_tokens=budget), 772 ) 773 per_step = [s.usage.input_tokens + s.usage.output_tokens for s in run.steps] 774 fig, ax = plt.subplots(figsize=(7, 4)) 775 ax.bar(range(len(per_step)), per_step, color="#7aa6d8", label="tokens sent this step") 776 ax.plot(range(len(per_step)), np.cumsum(per_step), "o-", color="#c0392b", label="running total") 777 ax.axhline(budget, ls="--", color="gray", label=f"budget ({budget:,})") 778 ax.set(xlabel="step", ylabel="tokens", title="A runaway agent: each call re-sends the whole history") 779 ax.legend() 780 figs["tokens_per_step"] = fig 781 782 # 2. Linear intuition vs. quadratic reality. 783 n = np.arange(1, 31) 784 t = 500 785 fig, ax = plt.subplots(figsize=(7, 4)) 786 ax.plot(n, t * n, label="if each step cost the same (linear)") 787 ax.plot(n, [cumulative_input_tokens(int(k), t) for k in n], label="re-sending history (quadratic)") 788 ax.set(xlabel="steps in the run", ylabel="total input tokens", title=f"Cumulative input tokens ({t} new tokens per step)") 789 ax.legend() 790 figs["quadratic_cost"] = fig 791 792 # 3. Sequential vs. concurrent tool execution. 793 latencies = np.random.default_rng(0).uniform(0.2, 1.2, 8) 794 k = np.arange(1, 9) 795 fig, ax = plt.subplots(figsize=(7, 4)) 796 ax.plot(k, [latencies[:i].sum() for i in k], "o-", label="sequential (sum of latencies)") 797 ax.plot(k, [latencies[:i].max() for i in k], "o-", label="concurrent (max latency)") 798 ax.set(xlabel="tool calls requested in one turn", ylabel="wall-clock seconds", title="Parallel tool calls: time is the slowest call, not the sum") 799 ax.legend() 800 figs["parallel_latency"] = fig 801 return figs 802 803 804def demo() -> None: 805 banner("1. Happy path: parallel tool calls, then an answer") 806 r = run_agent(ScriptedLLM(happy_policy), "How many PTO days does alice have left, and what rolls over?", DEMO_TOOLS, DEMO_TOOL_DEFS) 807 say(f"stop_cause={r.stop_cause}, steps={len(r.steps)}, tool calls={r.tool_call_count}") 808 say(f"Final answer: {r.final_text}") 809 say( 810 """ 811 Step 0 asked for two tools in one turn. They ran concurrently and both 812 results went back in a single user message, which is what the API expects. 813 """ 814 ) 815 816 banner("2. A model that loops: caught by loop detection") 817 r = run_agent(ScriptedLLM(looping_policy), "What's the PTO rollover rule?", DEMO_TOOLS, DEMO_TOOL_DEFS) 818 say(f"stop_cause={r.stop_cause}: {r.detail}") 819 takeaway("Detect the same tool with the same canonical arguments; stop after N repeats instead of paying for N more.") 820 821 banner("3. A model that burns budget: caught by the token budget") 822 r = run_agent( 823 ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS, 824 config=AgentConfig(max_steps=50, max_total_tokens=12_000), 825 ) 826 table( 827 ["step", "input tokens", "output tokens"], 828 [(s.index, s.usage.input_tokens, s.usage.output_tokens) for s in r.steps], 829 ) 830 say(f"stop_cause={r.stop_cause}: {r.detail}. Estimated cost ${r.cost_usd:.4f}.") 831 say( 832 """ 833 Input tokens per step climb steadily because every call re-sends the 834 whole conversation. Total input grows with the square of the step count. 835 """ 836 ) 837 838 banner("4. Recovering from a tool error") 839 r = run_agent(ScriptedLLM(recovering_policy), "How many PTO days will bob have on Oct 2?", DEMO_TOOLS, DEMO_TOOL_DEFS) 840 for s in r.steps: 841 for ex in s.tool_executions: 842 print(f" step {s.index}: {ex.name}({ex.input}) -> {'ERROR: ' if ex.is_error else ''}{ex.content}") 843 print() 844 say(f"stop_cause={r.stop_cause}. Final: {r.final_text}") 845 takeaway("Errors are results, not crashes. An actionable message lets the model fix its own call.") 846 847 banner("5. Hand-off to a human") 848 r = run_agent(ScriptedLLM(handoff_policy), "Please refund my $900 conference ticket.", DEMO_TOOLS, DEMO_TOOL_DEFS) 849 say(f"stop_cause={r.stop_cause}. Reason: {r.final_text}") 850 851 852if __name__ == "__main__": 853 demo()
474class ToolError(Exception): 475 """Raise from a tool to send an actionable `is_error` result back to the model. 476 477 The message is shown to the model verbatim, so write it for the model: 478 say what was wrong and what a valid call looks like. 479 """
Raise from a tool to send an actionable is_error result back to the model.
The message is shown to the model verbatim, so write it for the model: say what was wrong and what a valid call looks like.
482@dataclass 483class AgentConfig: 484 max_steps: int = 10 485 max_total_tokens: int = 50_000 486 max_cost_usd: float | None = None 487 # Stop when the same (tool, args) pair has been requested this many times. 488 loop_threshold: int = 3 489 max_parallel_tools: int = 4 490 max_tokens_per_call: int = 4096 491 # Dollar prices per million tokens. Defaults are Claude Opus 5 list prices 492 # at the time of writing; check the provider's pricing page. 493 price_in_per_mtok: float = 5.0 494 price_out_per_mtok: float = 25.0 495 496 def cost(self, usage: Usage) -> float: 497 return (usage.input_tokens * self.price_in_per_mtok + usage.output_tokens * self.price_out_per_mtok) / 1e6
500@dataclass 501class ToolExecution: 502 name: str 503 input: dict[str, Any] 504 content: str 505 is_error: bool
508@dataclass 509class Step: 510 """One model call plus the tools it triggered.""" 511 512 index: int 513 text: str 514 tool_executions: list[ToolExecution] 515 usage: Usage 516 stop_reason: str
One model call plus the tools it triggered.
519@dataclass 520class AgentResult: 521 final_text: str 522 stop_cause: StopCause 523 steps: list[Step] 524 usage: Usage 525 cost_usd: float 526 transcript: list[dict[str, Any]] = field(repr=False) 527 detail: str = "" # human-readable explanation of the stop cause 528 529 @property 530 def ok(self) -> bool: 531 return self.stop_cause == "completed" 532 533 @property 534 def tool_call_count(self) -> int: 535 return sum(len(s.tool_executions) for s in self.steps)
562def execute_tools(calls: list[ToolCall], tools: dict[str, Callable[..., Any]], max_workers: int = 4) -> list[ToolExecution]: 563 """Run tool calls concurrently. Results come back in the same order as `calls`. 564 565 Tools in one assistant turn are independent by construction (the model 566 asked for them together, before seeing any result), so running them in 567 parallel is safe and cuts wall-clock time to the slowest call. 568 """ 569 if len(calls) <= 1: 570 return [_execute_one(c, tools) for c in calls] 571 with ThreadPoolExecutor(max_workers=max_workers) as pool: 572 return list(pool.map(lambda c: _execute_one(c, tools), calls))
Run tool calls concurrently. Results come back in the same order as calls.
Tools in one assistant turn are independent by construction (the model asked for them together, before seeing any result), so running them in parallel is safe and cuts wall-clock time to the slowest call.
575def run_agent( 576 llm: LLM, 577 task: str, 578 tools: dict[str, Callable[..., Any]], 579 tool_defs: list[dict[str, Any]], 580 system: str = "You are a helpful assistant. Use the tools when needed.", 581 config: AgentConfig | None = None, 582) -> AgentResult: 583 """Run the think -> act -> observe loop until done or a control trips. 584 585 Args: 586 llm: anything implementing `LLM.complete` (ScriptedLLM or ClaudeLLM). 587 task: the user's request. 588 tools: tool name -> Python callable taking the tool input as kwargs. 589 tool_defs: Anthropic-format tool definitions sent to the model. 590 system: system prompt. 591 config: limits and prices. 592 593 Returns: 594 AgentResult with the final text, stop cause, per-step record, usage, 595 cost and the full transcript (for tracing and replay). 596 """ 597 config = config or AgentConfig() 598 messages: list[dict[str, Any]] = [{"role": "user", "content": task}] 599 steps: list[Step] = [] 600 usage = Usage() 601 seen: Counter[str] = Counter() 602 603 def finish(cause: StopCause, text: str = "", detail: str = "") -> AgentResult: 604 return AgentResult(text, cause, steps, usage, config.cost(usage), messages, detail) 605 606 for i in range(config.max_steps): 607 resp: LLMResponse = llm.complete( 608 system=system, messages=messages, tools=tool_defs, max_tokens=config.max_tokens_per_call 609 ) 610 usage += resp.usage 611 # Always append the full assistant content (it may carry thinking blocks 612 # that must be sent back unchanged with the real API). 613 messages.append({"role": "assistant", "content": resp.assistant_content}) 614 step = Step(i, resp.text, [], resp.usage, resp.stop_reason) 615 steps.append(step) 616 617 # --- Stop reasons that aren't "use a tool" ------------------------- 618 if resp.stop_reason == "refusal": 619 return finish("refusal", resp.text, "model declined the request") 620 if resp.stop_reason == "max_tokens": 621 # A truncated response may contain a half-formed tool call. Don't run it. 622 return finish("max_tokens", resp.text, "response truncated; raise max_tokens or ask for less output") 623 if resp.stop_reason == "end_turn" or not resp.tool_calls: 624 return finish("completed", resp.text) 625 626 # --- Budget: checked after every call, before spending more --------- 627 total = usage.input_tokens + usage.output_tokens 628 if total > config.max_total_tokens: 629 return finish("budget_exceeded", resp.text, f"{total:,} tokens > budget {config.max_total_tokens:,}") 630 if config.max_cost_usd is not None and config.cost(usage) > config.max_cost_usd: 631 return finish("budget_exceeded", resp.text, f"${config.cost(usage):.4f} > budget ${config.max_cost_usd:.4f}") 632 633 # --- Hand-off: an explicit, logged outcome -------------------------- 634 for call in resp.tool_calls: 635 if call.name == HANDOFF_TOOL: 636 return finish("handoff", call.input.get("reason", ""), "escalated to a human") 637 638 # --- Loop detection: same call requested too many times ------------- 639 for call in resp.tool_calls: 640 seen[_signature(call)] += 1 641 if seen[_signature(call)] >= config.loop_threshold: 642 return finish( 643 "loop_detected", resp.text, f"{call.name}({json.dumps(call.input)}) requested {seen[_signature(call)]} times" 644 ) 645 646 # --- Act: run all requested tools concurrently ---------------------- 647 executions = execute_tools(resp.tool_calls, tools, config.max_parallel_tools) 648 step.tool_executions = executions 649 # All results go back in ONE user message, matched by tool_use_id. 650 messages.append( 651 { 652 "role": "user", 653 "content": [ 654 tool_result_block(call.id, ex.content, ex.is_error) for call, ex in zip(resp.tool_calls, executions) 655 ], 656 } 657 ) 658 659 return finish("max_steps", steps[-1].text if steps else "", f"hit max_steps={config.max_steps}")
Run the think -> act -> observe loop until done or a control trips.
Arguments:
- llm: anything implementing
LLM.complete(ScriptedLLM or ClaudeLLM). - task: the user's request.
- tools: tool name -> Python callable taking the tool input as kwargs.
- tool_defs: Anthropic-format tool definitions sent to the model.
- system: system prompt.
- config: limits and prices.
Returns:
AgentResult with the final text, stop cause, per-step record, usage, cost and the full transcript (for tracing and replay).
687def search_kb(query: str) -> str: 688 """Tiny keyword search over the shared corpus (see primer.agents.rag for the real thing).""" 689 from primer.common.corpus import DOCS 690 from primer.common.text import tokenize 691 692 q = set(tokenize(query)) 693 scored = sorted(DOCS, key=lambda d: -len(q & set(tokenize(d.title + " " + d.text)))) 694 return "\n".join(f"[{d.id}] {d.title}: {d.text}" for d in scored[:2])
Tiny keyword search over the shared corpus (see primer.agents.rag for the real thing).
697def get_pto_balance(employee: str, as_of: str) -> str: 698 import re 699 700 if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", as_of): 701 # Actionable: says what's wrong AND what a valid value looks like. 702 raise ToolError(f"as_of must be a date in YYYY-MM-DD format (e.g. 2026-10-02), got {as_of!r}.") 703 if employee not in PTO_BALANCES: 704 raise ToolError(f"Unknown employee {employee!r}. Known usernames: {sorted(PTO_BALANCES)}.") 705 return json.dumps({"employee": employee, "as_of": as_of, "days_left": PTO_BALANCES[employee]})
712def happy_policy(system, messages, tools): 713 """Turn 1: two independent lookups in parallel. Turn 2: answer from the results.""" 714 results = tool_results(messages) 715 if not results: 716 return ( 717 "I'll check the PTO policy and alice's balance at the same time.", 718 [ToolCall("", "search_kb", {"query": "PTO rollover"}), ToolCall("", "get_pto_balance", {"employee": "alice", "as_of": "2026-09-25"})], 719 ) 720 balance = json.loads(results[-1]["content"])["days_left"] 721 return f"Alice has {balance} PTO days left. Up to 5 unused days roll over to next year [hr-001]."
Turn 1: two independent lookups in parallel. Turn 2: answer from the results.
724def looping_policy(system, messages, tools): 725 """A confused model that keeps re-running the same search.""" 726 return ToolCall("", "search_kb", {"query": "PTO rollover"})
A confused model that keeps re-running the same search.
729def budget_burner_policy(system, messages, tools): 730 """A model that keeps searching with new queries: no loop, just ever-growing context.""" 731 n = len(tool_results(messages)) 732 return ToolCall("", "search_kb", {"query": f"PTO policy details page {n}"})
A model that keeps searching with new queries: no loop, just ever-growing context.
735def recovering_policy(system, messages, tools): 736 """First call has a bad date; the model reads the actionable error and fixes it.""" 737 results = tool_results(messages) 738 if not results: 739 return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "next friday"}) 740 last = results[-1] 741 if last.get("is_error"): 742 return ToolCall("", "get_pto_balance", {"employee": "bob", "as_of": "2026-10-02"}) 743 return f"Bob will have {json.loads(last['content'])['days_left']} PTO days left on 2026-10-02."
First call has a bad date; the model reads the actionable error and fixes it.
752def cumulative_input_tokens(n_steps: int, tokens_per_step: int) -> int: 753 """Total input tokens for an n-step run when step k re-sends k·t tokens of history.""" 754 return tokens_per_step * n_steps * (n_steps + 1) // 2
Total input tokens for an n-step run when step k re-sends k·t tokens of history.
757def figures() -> dict: 758 """Plots computed from this lesson's own code. matplotlib is imported here 759 so the lesson itself needs only NumPy.""" 760 import matplotlib 761 762 matplotlib.use("Agg") 763 import matplotlib.pyplot as plt 764 import numpy as np 765 766 figs = {} 767 768 # 1. The budget-burner run, step by step. 769 budget = 12_000 770 run = run_agent( 771 ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS, 772 config=AgentConfig(max_steps=50, max_total_tokens=budget), 773 ) 774 per_step = [s.usage.input_tokens + s.usage.output_tokens for s in run.steps] 775 fig, ax = plt.subplots(figsize=(7, 4)) 776 ax.bar(range(len(per_step)), per_step, color="#7aa6d8", label="tokens sent this step") 777 ax.plot(range(len(per_step)), np.cumsum(per_step), "o-", color="#c0392b", label="running total") 778 ax.axhline(budget, ls="--", color="gray", label=f"budget ({budget:,})") 779 ax.set(xlabel="step", ylabel="tokens", title="A runaway agent: each call re-sends the whole history") 780 ax.legend() 781 figs["tokens_per_step"] = fig 782 783 # 2. Linear intuition vs. quadratic reality. 784 n = np.arange(1, 31) 785 t = 500 786 fig, ax = plt.subplots(figsize=(7, 4)) 787 ax.plot(n, t * n, label="if each step cost the same (linear)") 788 ax.plot(n, [cumulative_input_tokens(int(k), t) for k in n], label="re-sending history (quadratic)") 789 ax.set(xlabel="steps in the run", ylabel="total input tokens", title=f"Cumulative input tokens ({t} new tokens per step)") 790 ax.legend() 791 figs["quadratic_cost"] = fig 792 793 # 3. Sequential vs. concurrent tool execution. 794 latencies = np.random.default_rng(0).uniform(0.2, 1.2, 8) 795 k = np.arange(1, 9) 796 fig, ax = plt.subplots(figsize=(7, 4)) 797 ax.plot(k, [latencies[:i].sum() for i in k], "o-", label="sequential (sum of latencies)") 798 ax.plot(k, [latencies[:i].max() for i in k], "o-", label="concurrent (max latency)") 799 ax.set(xlabel="tool calls requested in one turn", ylabel="wall-clock seconds", title="Parallel tool calls: time is the slowest call, not the sum") 800 ax.legend() 801 figs["parallel_latency"] = fig 802 return figs
Plots computed from this lesson's own code. matplotlib is imported here so the lesson itself needs only NumPy.
805def demo() -> None: 806 banner("1. Happy path: parallel tool calls, then an answer") 807 r = run_agent(ScriptedLLM(happy_policy), "How many PTO days does alice have left, and what rolls over?", DEMO_TOOLS, DEMO_TOOL_DEFS) 808 say(f"stop_cause={r.stop_cause}, steps={len(r.steps)}, tool calls={r.tool_call_count}") 809 say(f"Final answer: {r.final_text}") 810 say( 811 """ 812 Step 0 asked for two tools in one turn. They ran concurrently and both 813 results went back in a single user message, which is what the API expects. 814 """ 815 ) 816 817 banner("2. A model that loops: caught by loop detection") 818 r = run_agent(ScriptedLLM(looping_policy), "What's the PTO rollover rule?", DEMO_TOOLS, DEMO_TOOL_DEFS) 819 say(f"stop_cause={r.stop_cause}: {r.detail}") 820 takeaway("Detect the same tool with the same canonical arguments; stop after N repeats instead of paying for N more.") 821 822 banner("3. A model that burns budget: caught by the token budget") 823 r = run_agent( 824 ScriptedLLM(budget_burner_policy), "Tell me everything about PTO.", DEMO_TOOLS, DEMO_TOOL_DEFS, 825 config=AgentConfig(max_steps=50, max_total_tokens=12_000), 826 ) 827 table( 828 ["step", "input tokens", "output tokens"], 829 [(s.index, s.usage.input_tokens, s.usage.output_tokens) for s in r.steps], 830 ) 831 say(f"stop_cause={r.stop_cause}: {r.detail}. Estimated cost ${r.cost_usd:.4f}.") 832 say( 833 """ 834 Input tokens per step climb steadily because every call re-sends the 835 whole conversation. Total input grows with the square of the step count. 836 """ 837 ) 838 839 banner("4. Recovering from a tool error") 840 r = run_agent(ScriptedLLM(recovering_policy), "How many PTO days will bob have on Oct 2?", DEMO_TOOLS, DEMO_TOOL_DEFS) 841 for s in r.steps: 842 for ex in s.tool_executions: 843 print(f" step {s.index}: {ex.name}({ex.input}) -> {'ERROR: ' if ex.is_error else ''}{ex.content}") 844 print() 845 say(f"stop_cause={r.stop_cause}. Final: {r.final_text}") 846 takeaway("Errors are results, not crashes. An actionable message lets the model fix its own call.") 847 848 banner("5. Hand-off to a human") 849 r = run_agent(ScriptedLLM(handoff_policy), "Please refund my $900 conference ticket.", DEMO_TOOLS, DEMO_TOOL_DEFS) 850 say(f"stop_cause={r.stop_cause}. Reason: {r.final_text}")