primer.agents.llm

Talking to a model: messages, and what tool calling really is

Run: python -m primer.agents.llm

New to the notation? primer.notation explains every symbol used here from zero. This is the first lesson of the agents part: it builds on how a model produces text (primer.ml.inference), and every later agent lesson builds on it.

Level 1: The practitioner's guide

In one sentence. Talking to a model means sending it the whole conversation so far as a list of messages and getting one message back, and tool calling is the case where that message is a structured request ("run get_weather with {"city": "Paris"}") that your code may carry out and answer.

When you need it. Every system built on a language model does this, from a one-shot classifier to a multi-agent research team, so the message format is not optional. Tool calling is. You need it the moment the model must reach past its own weights: look something up, run a query, compute a number, send an email, change a record. You don't need it when the answer is words the model already knows, and you don't need it when one structured answer is enough (a label, a JSON row): that is structured output (primer.ml.structured_output), which is a tool call without the "run it and come back" step. The tell: if your code reads the model's prose to decide which function to call, or scrapes a city name out of a sentence, you need tool calling. The model was trained to hand you that request as data.

Two facts about the exchange decide most of what follows, and both come from this lesson's traced example. First, the model executes nothing: the reply to "What's the weather in Paris?" is a tool_use block asking for get_weather, and the 18°C comes from your code running the function and sending the result back. Second, the model keeps no state between calls, so every call carries the whole history again. In the lesson's ten-call loop, that turns a conversation 7,000 tokens long into 42,500 input tokens billed.

Your options. Five ways to get an action out of a model, from the cheapest to the most certain:

Option What it does What it guarantees What it costs Where it lives
A scripted stand-in A function you write plays the model, replying by rule The exact reply you scripted, every run, offline Nothing per call; proves nothing about a real model Your tests and demos
A local open model, text only Weights on your own machine answer in plain text Privacy: nothing leaves the box; no per-token bill Hardware, smaller models, and your own parsing if you need actions Your machine (Ollama and similar)
Text you parse Ask the model to write the action in prose, then match it with a regular expression Nothing; it works until the wording drifts A parser you maintain and a retry loop Your code
Native tool calling The model replies with a typed tool_use block: a name, an id, and the arguments as a parsed JSON object A request in your schema's shape; with strict mode, exactly your schema Tokens for the tool definitions, a tool-use system prompt the API adds (286 tokens on Claude Opus 5, per the pricing page), and one extra round trip per call The model server produces it; your code runs it
Server-side tools The provider runs the tool (web search, web fetch, code execution) and returns the result in the same reply The result arrives with no handler code on your side Per-use fees on some tools (Claude's web search is \$10 per 1,000 searches) and no control over execution The model server

How to choose. Start from what has to happen, and who is allowed to make it happen.

  • Testing the code around the model (a loop, a budget, an approval step): the scripted stand-in. It reproduces a model that loops, calls the wrong tool or follows an injected instruction on cue, which a real model rarely does when you want it to.
  • Private data, no budget, or a laptop on a plane: a local open model. Keep it to text unless the model and server both support tool calling (Ollama does, for models trained for it; this lesson's local adapter is text only).
  • Anything a program acts on: native tool calling, never parsed prose. Turn on strict mode when the arguments feed straight into code.
  • Web search or sandboxed code that you would otherwise host yourself: a server-side tool, if its fees and its lack of a hook for your own checks suit you.
  • Whatever you pick, your code stays the only thing that acts. Validate the arguments, check permissions, then run the tool; the model only asks.

What it costs. Input tokens are billed on the whole history, every call, so the cost of a loop grows with the square of its length: the lesson's ten calls bill 42,500 tokens, six times the 7,000-token conversation they end with, and the tenth call alone re-sends 6,500. At Claude Opus 5's list price of \$5 per million input tokens (the pricing page, fetched for this guide) that is about 21 cents per ten-step task before any output; on Haiku 4.5 at \$1 per million, about 4 cents. Prompt caching sells re-read tokens at a tenth of the price, which is why it is the first lever on any loop (primer.agents.cost). Tool definitions ride along on every call too, so a long tool list is a standing charge. Latency is one network round trip per model call plus the tool's own time, and a tool call always means at least two model calls.

What breaks.

  • Dropping the assistant turn. Append the exact content blocks the model returned, not just their text: the ids in them are how the next call matches results to requests, and real APIs can include blocks (thinking) that must go back unchanged.
  • Unmatched ids. A tool_result whose tool_use_id matches no request answers nothing. Return one result per request, and all of them in a single user message when the model asked for several.
  • Swallowing failures. A tool that throws and returns nothing leaves the model waiting. Send a result flagged as an error with a message the model can act on, so it can retry or ask.
  • Trusting the request. Arguments arrive as a parsed object in the shape of your schema, not as safe values. Validate them and check permissions before running anything; the model is never a security boundary.
  • Ignoring stop_reason. tool_use means run tools and call again; end_turn means done; max_tokens means the reply was cut off and its JSON may be incomplete; refusal means the model declined. A loop that checks only for text gets all four wrong.
  • Missing arguments guessed. Asked for the weather with no city, a model may invent one rather than ask (Claude's tool-use docs call the asking behaviour model-dependent and not guaranteed). Make optional what is optional, and validate the rest.

In the wild. Claude's Messages API is the shape this lesson uses directly: tool_use and tool_result blocks, stop_reason, parallel tool calls, strict: true for schema-exact arguments, and server tools (web search, web fetch, code execution) alongside the client tools you run. Its SDKs add a tool runner that drives the request, run, reply loop for you. OpenAI's function calling is the same five-step round trip in different field names, with its own strict mode; the guide recommends turning it on. Ollama exposes tool calling for local models, where again the model returns the call and your code runs it. Every agent framework, from the smallest loop to a multi-agent system, is built on this exchange, and the papers behind it (Toolformer, ReAct) are in this lesson's papers section.

Go deeper. Level 2 traces the four messages of one tool call by hand, lays out the message format field by field, shows the one interface that a scripted model and a real one both implement, and derives the quadratic cost of a loop with a formula you can rerun. If you only needed to choose, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

Every applied-AI system in this primer, from a one-shot classifier to a multi-agent research team, talks to the model the same way: it sends a list of messages and gets one message back. This lesson shows that exchange exactly, and then the one feature that turns a chatbot into an agent: tool calling.

The everyday picture

Think of the model as a brilliant consultant who works by post. You mail a letter with your whole conversation so far (the consultant keeps no notes between letters) and get one letter back. The reply is either an answer, or a filled-in request form: "please look up the weather in Paris and send me the result." The consultant never picks up the phone themselves. You decide whether to run the request, run it, and mail back the result, together with the whole conversation again.

Two facts fall straight out of this picture, and they explain most of the behaviour of real systems:

  1. The model executes nothing. A tool call is a request. Your code is the only thing that acts, so your code is where safety lives.
  2. Every call resends everything. The model has no memory between calls, so each request carries the full history. Long conversations cost more on every single turn.

A tiny worked example: one tool call, traced

A user asks "What's the weather in Paris?" and the application offers one tool, get_weather. Four messages later, the user has an answer:

# Role Content Who produced it
1 user "What's the weather in Paris?" the user
2 assistant tool_use id=toolu_1, name=get_weather, input={"city": "Paris"} the model (1st call)
3 user tool_result for toolu_1: "18°C, sunny" your code, after running the tool
4 assistant "It's 18°C and sunny in Paris." the model (2nd call)

Message 3 has role user even though no human typed it: tool results always travel back in the user's turn. The id in message 2 and the tool_use_id in message 3 match, which is how the model knows which request a result answers when it asked for several at once. trace_tool_call_round_trip() produces exactly this transcript.

sequenceDiagram participant U as User participant A as Your code participant M as Model participant T as get_weather tool U->>A: What's the weather in Paris? A->>M: messages [1] + tool definitions M-->>A: tool_use get_weather {city: Paris} (stop_reason: tool_use) A->>A: validate the arguments, check permissions A->>T: get_weather("Paris") T-->>A: 18°C, sunny A->>M: messages [1, 2, 3 = tool_result] M-->>A: It's 18°C and sunny in Paris. (stop_reason: end_turn) A-->>U: It's 18°C and sunny in Paris.

Reading it: time runs downwards. Solid arrows are requests and dashed arrows are replies. Notice that the model never talks to the tool: every arrow into get_weather starts at your code, and the self-arrow on your code ("validate, check permissions") is where production systems put their guardrails. Notice too that the second call to the model carries messages 1 to 3, not just the new result: the model is stateless, so the history travels every time.

In code: ToolCall holds the request in message 2, and tool_result_block builds the reply in message 3, carrying the matching id.

The message format

Messages use the Anthropic Messages API shape directly, so what you learn here maps one-to-one onto production code:

{"role": "user", "content": "What's our PTO policy?"}
{"role": "assistant", "content": [
    {"type": "text", "text": "Let me look that up."},
    {"type": "tool_use", "id": "toolu_1", "name": "search_kb", "input": {"query": "PTO"}},
]}
{"role": "user", "content": [
    {"type": "tool_result", "tool_use_id": "toolu_1", "content": "hr-001: ..."},
]}
Field Meaning
role user (the human, or your code returning tool results) or assistant (the model)
content a plain string, or a list of content blocks
text block ordinary words
tool_use block the model asking for a tool: id, name, and input (a JSON object matching the tool's schema)
tool_result block your answer to one tool_use, matched by tool_use_id; set is_error: true when the tool failed
stop_reason why the reply ended: end_turn (finished), tool_use (wants a tool), max_tokens (ran out of room), refusal (declined)

A tool definition is a name, a description and a JSON Schema for the input. The model chooses tools by reading those descriptions, so writing them well is prompt engineering (see primer.agents.tools).

In code: LLMResponse is one reply, normalized: its text, its ToolCall list, its stop reason, its Usage and the exact content blocks to append as the assistant turn. content_blocks, last_user_text, tool_results and tool_calls_so_far read a conversation in this format.

Two implementations of one interface

Every agent lesson is written against one small interface, LLM, with one method, LLM.complete(system=..., messages=..., tools=...). Two classes implement it:

flowchart LR L["LLM interface<br/>complete(system, messages, tools)"] L --> S["ScriptedLLM<br/>a Python function decides the reply<br/>offline, instant, deterministic"] L --> C["ClaudeLLM<br/>the real model via the anthropic SDK<br/>needs an API key"] S --> D["demos and tests<br/>reproduce any failure on purpose"] C --> P["production<br/>same agent code, real behaviour"]

Reading it: the agent code on the right-hand side never knows which box it's talking to. ScriptedLLM lets every lesson run offline and deterministically, and lets tests stage a model that loops, calls the wrong tool or follows an injected instruction, which is hard to get a real model to do on cue. Swapping in ClaudeLLM runs the identical loop against the real thing.

In code: ClaudeLLM.complete builds its request with claude_request and turns the API's reply into an LLMResponse. OllamaLLM is a third implementation for a local open model (text only), using ollama_request and parse_ollama_reply.

Cost: why every call pays for the whole conversation

Because the model is stateless, input tokens are billed per call on the entire history. In an agent loop with $n$ calls, where each call adds about $t$ new tokens to a history that started at $h_0$ tokens, the total input billed is:

Level 3: the formula and its symbols

$$ \text{input tokens} \approx \sum_{i=1}^{n} \left(h_0 + (i-1)\,t\right) = n\,h_0 + t\,\frac{n(n-1)}{2} $$

Symbols

Symbol Meaning here In the example
$n$ number of model calls in the loop 10
$h_0$ tokens in the first request (system prompt, tools, question) 2,000
$t$ tokens each step adds (the model's tool call plus the tool's result) 500
$i$ which call we're on, 1 to $n$
$\sum_{i=1}^{n}$ add up the cost of every call
$\frac{n(n-1)}{2}$ 0 + 1 + … + (n − 1): how many "steps of history" pile up in total 45

In words: "each call pays for the starting prompt plus everything added so far, so the history term grows with the square of the number of steps."

With the numbers: 10 × 2,000 + 500 × 45 = 20,000 + 22,500 = 42,500 input tokens for a task whose final conversation is only 7,000 tokens long: the tenth call re-sends 6,500 tokens, and its own 500-token step brings the history to 7,000.

Level 3: in Python

In Python:

n, h_0, t = 10, 2000, 500
# call i re-sends h_0 and i - 1 steps
calls = [h_0 + (i - 1) * t for i in range(1, n + 1)]
calls[0], calls[-1]  # → (2000, 6500)
# Σ over every call
sum(calls)  # → 42500
# the shortcut on the right agrees
n * h_0 + t * n * (n - 1) // 2  # → 42500

Ten calls bill 42,500 input tokens in total, six times the 7,000-token conversation they end with, because every call re-sends the history

Reading it: the bars are the input tokens billed on each call. They grow by the same amount every step, because each call re-sends the whole history. The line is the running total, and it curves upwards (quadratic growth). The dashed line is what you might naively expect: the size of the final conversation, billed once. The gap between the dashed line and the curve is why prompt caching, trimming tool outputs and keeping loops short are the main cost levers (primer.agents.cost, primer.agents.context).

In code: loop_input_tokens lists the tokens billed on each call of such a loop. estimate_tokens is the four-characters-per-token rule of thumb, and conversation_chars measures everything a call resends, which is how ScriptedLLM gives its fake replies realistic Usage.

In 20 seconds

  • A model call sends the full list of messages and returns one message; the model keeps no state between calls.
  • A tool call is a structured request (tool_use). Your code validates it, runs it, and returns a tool_result with the matching id. The model executes nothing.
  • stop_reason tells you what to do next: tool_use means run tools and call again, end_turn means done.
  • Every call re-sends the whole history, so input cost grows quadratically with the number of steps in a loop.

Self-test questions

What exactly happens when a model "uses a tool"? You send tool definitions (name, description, JSON Schema) with the messages. The model replies with a tool_use block and stop_reason: tool_use. Your code validates the arguments, runs the function, and sends a new request with the history plus a tool_result block carrying the same id. The model then answers or asks for another tool.

Why is the model never a security boundary for tool use? It only produces requests. Everything that actually happens goes through your code, so validation, permissions, approvals and rate limits all belong there, and a manipulated model can only ever ask.

The model asks for three tools in one reply. How do you send the results? Run them (concurrently if independent) and return all three tool_result blocks in a single user message, each with its own tool_use_id. For one that failed, return a tool_result with is_error: true and an actionable message rather than dropping it.

A 20-step agent loop costs far more than 20 times a single call. Why? Each call re-sends the full, growing history, so the input billed is a sum that grows with the square of the number of steps. Caching the stable prefix and trimming tool outputs attack exactly this.

Why test agents against a scripted model at all? Real models are non-deterministic and rarely misbehave on cue. A scripted model reproduces loops, bad arguments and injected instructions exactly, so the code that must handle them can be tested every time.

The papers behind this lesson

Further reading

on GitHub
  1r"""
  2# Talking to a model: messages, and what tool calling really is
  3
  4Run: `python -m primer.agents.llm`
  5
  6New to the notation? `primer.notation` explains every symbol used here from
  7zero. This is the first lesson of the agents part: it builds on how a model
  8produces text (`primer.ml.inference`), and every later agent lesson builds on
  9it.
 10
 11## Level 1: The practitioner's guide
 12
 13**In one sentence.** Talking to a model means sending it the whole
 14conversation so far as a list of messages and getting one message back, and
 15tool calling is the case where that message is a structured request ("run
 16`get_weather` with `{"city": "Paris"}`") that your code may carry out and
 17answer.
 18
 19**When you need it.** Every system built on a language model does this,
 20from a one-shot classifier to a multi-agent research team, so the message
 21format is not optional. Tool calling is. You need it the moment the model
 22must reach past its own weights: look something up, run a query, compute a
 23number, send an email, change a record. You don't need it when the answer is
 24words the model already knows, and you don't need it when one structured
 25answer is enough (a label, a JSON row): that is structured output
 26(`primer.ml.structured_output`), which is a tool call without the "run it and
 27come back" step. The tell: if your code reads the model's prose to decide
 28which function to call, or scrapes a city name out of a sentence, you need
 29tool calling. The model was trained to hand you that request as data.
 30
 31Two facts about the exchange decide most of what follows, and both come from
 32this lesson's traced example. First, the model executes nothing: the reply
 33to "What's the weather in Paris?" is a `tool_use` block asking for
 34`get_weather`, and the 18°C comes from your code running the function and
 35sending the result back. Second, the model keeps no state between calls, so
 36every call carries the whole history again. In the lesson's ten-call loop,
 37that turns a conversation 7,000 tokens long into 42,500 input tokens billed.
 38
 39**Your options.** Five ways to get an action out of a model, from the
 40cheapest to the most certain:
 41
 42| Option | What it does | What it guarantees | What it costs | Where it lives |
 43|---|---|---|---|---|
 44| A scripted stand-in | A function you write plays the model, replying by rule | The exact reply you scripted, every run, offline | Nothing per call; proves nothing about a real model | Your tests and demos |
 45| A local open model, text only | Weights on your own machine answer in plain text | Privacy: nothing leaves the box; no per-token bill | Hardware, smaller models, and your own parsing if you need actions | Your machine (Ollama and similar) |
 46| Text you parse | Ask the model to write the action in prose, then match it with a regular expression | Nothing; it works until the wording drifts | A parser you maintain and a retry loop | Your code |
 47| Native tool calling | The model replies with a typed `tool_use` block: a name, an id, and the arguments as a parsed JSON object | A request in your schema's shape; with strict mode, exactly your schema | Tokens for the tool definitions, a tool-use system prompt the API adds (286 tokens on Claude Opus 5, per the pricing page), and one extra round trip per call | The model server produces it; your code runs it |
 48| Server-side tools | The provider runs the tool (web search, web fetch, code execution) and returns the result in the same reply | The result arrives with no handler code on your side | Per-use fees on some tools (Claude's web search is \$10 per 1,000 searches) and no control over execution | The model server |
 49
 50**How to choose.** Start from what has to happen, and who is allowed to make
 51it happen.
 52
 53- Testing the code around the model (a loop, a budget, an approval step):
 54  the scripted stand-in. It reproduces a model that loops, calls the wrong
 55  tool or follows an injected instruction on cue, which a real model rarely
 56  does when you want it to.
 57- Private data, no budget, or a laptop on a plane: a local open model. Keep
 58  it to text unless the model and server both support tool calling (Ollama
 59  does, for models trained for it; this lesson's local adapter is text only).
 60- Anything a program acts on: native tool calling, never parsed prose. Turn
 61  on strict mode when the arguments feed straight into code.
 62- Web search or sandboxed code that you would otherwise host yourself: a
 63  server-side tool, if its fees and its lack of a hook for your own checks
 64  suit you.
 65- Whatever you pick, your code stays the only thing that acts. Validate the
 66  arguments, check permissions, then run the tool; the model only asks.
 67
 68**What it costs.** Input tokens are billed on the whole history, every call,
 69so the cost of a loop grows with the square of its length: the lesson's ten
 70calls bill 42,500 tokens, six times the 7,000-token conversation they end
 71with, and the tenth call alone re-sends 6,500. At Claude Opus 5's list price
 72of \$5 per million input tokens (the pricing page, fetched for this guide)
 73that is about 21 cents per ten-step task before any output; on Haiku 4.5 at
 74\$1 per million, about 4 cents. Prompt caching sells re-read tokens at a
 75tenth of the price, which is why it is the first lever on any loop
 76(`primer.agents.cost`). Tool definitions ride along on every call too, so a
 77long tool list is a standing charge. Latency is one network round trip per
 78model call plus the tool's own time, and a tool call always means at least
 79two model calls.
 80
 81**What breaks.**
 82
 83- **Dropping the assistant turn.** Append the exact content blocks the model
 84  returned, not just their text: the ids in them are how the next call
 85  matches results to requests, and real APIs can include blocks (thinking)
 86  that must go back unchanged.
 87- **Unmatched ids.** A `tool_result` whose `tool_use_id` matches no request
 88  answers nothing. Return one result per request, and all of them in a single
 89  user message when the model asked for several.
 90- **Swallowing failures.** A tool that throws and returns nothing leaves the
 91  model waiting. Send a result flagged as an error with a message the model
 92  can act on, so it can retry or ask.
 93- **Trusting the request.** Arguments arrive as a parsed object in the
 94  shape of your schema, not as safe values. Validate them and check
 95  permissions before running anything; the model is never a security
 96  boundary.
 97- **Ignoring `stop_reason`.** `tool_use` means run tools and call again;
 98  `end_turn` means done; `max_tokens` means the reply was cut off and its
 99  JSON may be incomplete; `refusal` means the model declined. A loop that
100  checks only for text gets all four wrong.
101- **Missing arguments guessed.** Asked for the weather with no city, a model
102  may invent one rather than ask (Claude's tool-use docs call the asking
103  behaviour model-dependent and not guaranteed). Make optional what is optional, and validate the
104  rest.
105
106**In the wild.** Claude's Messages API is the shape this lesson uses
107directly: `tool_use` and `tool_result` blocks, `stop_reason`, parallel tool
108calls, `strict: true` for schema-exact arguments, and server tools (web
109search, web fetch, code execution) alongside the client tools you run. Its
110SDKs add a tool runner that drives the request, run, reply loop for you.
111OpenAI's function calling is the same five-step round trip in different
112field names, with its own strict mode; the guide recommends turning it on.
113Ollama exposes tool calling for local models, where again the model returns
114the call and your code runs it. Every agent framework, from the smallest
115loop to a multi-agent system, is built on this exchange, and the papers
116behind it (Toolformer, ReAct) are in this lesson's papers section.
117
118**Go deeper.** Level 2 traces the four messages of one tool call by hand,
119lays out the message format field by field, shows the one interface that a
120scripted model and a real one both implement, and derives the quadratic
121cost of a loop with a formula you can rerun. If you only needed to choose,
122you are done.
123
124## Level 2: How it works, from scratch
125
126Every applied-AI system in this primer, from a one-shot classifier to a
127multi-agent research team, talks to the model the same way: it sends a list
128of **messages** and gets one message back. This lesson shows that exchange
129exactly, and then the one feature that turns a chatbot into an agent: **tool
130calling**.
131
132## The everyday picture
133
134Think of the model as a brilliant consultant who works by post. You mail a
135letter with your whole conversation so far (the consultant keeps no notes
136between letters) and get one letter back. The reply is either an answer, or
137a **filled-in request form**: "please look up the weather in Paris and send me
138the result." The consultant never picks up the phone themselves. *You* decide
139whether to run the request, run it, and mail back the result, together with
140the whole conversation again.
141
142Two facts fall straight out of this picture, and they explain most of the
143behaviour of real systems:
144
1451. **The model executes nothing.** A tool call is a request. Your code is
146   the only thing that acts, so your code is where safety lives.
1472. **Every call resends everything.** The model has no memory between calls,
148   so each request carries the full history. Long conversations cost more on
149   every single turn.
150
151## A tiny worked example: one tool call, traced
152
153A user asks "What's the weather in Paris?" and the application offers one
154tool, `get_weather`. Four messages later, the user has an answer:
155
156| # | Role | Content | Who produced it |
157|---|---|---|---|
158| 1 | user | "What's the weather in Paris?" | the user |
159| 2 | assistant | `tool_use` id=`toolu_1`, name=`get_weather`, input=`{"city": "Paris"}` | the model (1st call) |
160| 3 | user | `tool_result` for `toolu_1`: "18°C, sunny" | **your code**, after running the tool |
161| 4 | assistant | "It's 18°C and sunny in Paris." | the model (2nd call) |
162
163Message 3 has role `user` even though no human typed it: tool results always
164travel back in the user's turn. The `id` in message 2 and the `tool_use_id`
165in message 3 match, which is how the model knows which request a result
166answers when it asked for several at once. `trace_tool_call_round_trip()`
167produces exactly this transcript.
168
169```mermaid
170sequenceDiagram
171  participant U as User
172  participant A as Your code
173  participant M as Model
174  participant T as get_weather tool
175  U->>A: What's the weather in Paris?
176  A->>M: messages [1] + tool definitions
177  M-->>A: tool_use get_weather {city: Paris}  (stop_reason: tool_use)
178  A->>A: validate the arguments, check permissions
179  A->>T: get_weather("Paris")
180  T-->>A: 18°C, sunny
181  A->>M: messages [1, 2, 3 = tool_result]
182  M-->>A: It's 18°C and sunny in Paris.  (stop_reason: end_turn)
183  A-->>U: It's 18°C and sunny in Paris.
184```
185
186**Reading it:** time runs downwards. Solid arrows are requests and dashed
187arrows are replies. Notice that the model never talks to the tool: every
188arrow into `get_weather` starts at *your code*, and the self-arrow on your
189code ("validate, check permissions") is where production systems put their
190guardrails. Notice too that the second call to the model carries messages
1911 to 3, not just the new result: the model is stateless, so the history
192travels every time.
193
194**In code:** `ToolCall` holds the request in message 2, and
195`tool_result_block` builds the reply in message 3, carrying the matching id.
196
197## The message format
198
199Messages use the Anthropic Messages API shape directly, so what you learn
200here maps one-to-one onto production code:
201
202```python
203{"role": "user", "content": "What's our PTO policy?"}
204{"role": "assistant", "content": [
205    {"type": "text", "text": "Let me look that up."},
206    {"type": "tool_use", "id": "toolu_1", "name": "search_kb", "input": {"query": "PTO"}},
207]}
208{"role": "user", "content": [
209    {"type": "tool_result", "tool_use_id": "toolu_1", "content": "hr-001: ..."},
210]}
211```
212
213| Field | Meaning |
214|---|---|
215| `role` | `user` (the human, or your code returning tool results) or `assistant` (the model) |
216| `content` | a plain string, or a list of **content blocks** |
217| `text` block | ordinary words |
218| `tool_use` block | the model asking for a tool: `id`, `name`, and `input` (a JSON object matching the tool's schema) |
219| `tool_result` block | your answer to one `tool_use`, matched by `tool_use_id`; set `is_error: true` when the tool failed |
220| `stop_reason` | why the reply ended: `end_turn` (finished), `tool_use` (wants a tool), `max_tokens` (ran out of room), `refusal` (declined) |
221
222A **tool definition** is a name, a description and a JSON Schema for the
223input. The model chooses tools by reading those descriptions, so writing
224them well is prompt engineering (see `primer.agents.tools`).
225
226**In code:** `LLMResponse` is one reply, normalized: its text, its
227`ToolCall` list, its stop reason, its `Usage` and the exact content blocks
228to append as the assistant turn. `content_blocks`, `last_user_text`,
229`tool_results` and `tool_calls_so_far` read a conversation in this format.
230
231## Two implementations of one interface
232
233Every agent lesson is written against one small interface, `LLM`, with one
234method, `LLM.complete(system=..., messages=..., tools=...)`. Two classes
235implement it:
236
237```mermaid
238flowchart LR
239  L["LLM interface<br/>complete(system, messages, tools)"]
240  L --> S["ScriptedLLM<br/>a Python function decides the reply<br/>offline, instant, deterministic"]
241  L --> C["ClaudeLLM<br/>the real model via the anthropic SDK<br/>needs an API key"]
242  S --> D["demos and tests<br/>reproduce any failure on purpose"]
243  C --> P["production<br/>same agent code, real behaviour"]
244```
245
246**Reading it:** the agent code on the right-hand side never knows which
247box it's talking to. `ScriptedLLM` lets every lesson run offline and
248deterministically, and lets tests stage a model that loops, calls the wrong
249tool or follows an injected instruction, which is hard to get a real model
250to do on cue. Swapping in `ClaudeLLM` runs the identical loop against the
251real thing.
252
253**In code:** `ClaudeLLM.complete` builds its request with `claude_request`
254and turns the API's reply into an `LLMResponse`. `OllamaLLM` is a third
255implementation for a local open model (text only), using `ollama_request`
256and `parse_ollama_reply`.
257
258## Cost: why every call pays for the whole conversation
259
260Because the model is stateless, input tokens are billed per call on the
261entire history. In an agent loop with $n$ calls, where each call adds about
262$t$ new tokens to a history that started at $h_0$ tokens, the total input
263billed is:
264
265$$
266\text{input tokens} \approx \sum_{i=1}^{n} \left(h_0 + (i-1)\,t\right) = n\,h_0 + t\,\frac{n(n-1)}{2}
267$$
268
269**Symbols**
270
271| Symbol | Meaning here | In the example |
272|---|---|---|
273| $n$ | number of model calls in the loop | 10 |
274| $h_0$ | tokens in the first request (system prompt, tools, question) | 2,000 |
275| $t$ | tokens each step adds (the model's tool call plus the tool's result) | 500 |
276| $i$ | which call we're on, 1 to $n$ | |
277| $\sum_{i=1}^{n}$ | add up the cost of every call | |
278| $\frac{n(n-1)}{2}$ | 0 + 1 + … + (n − 1): how many "steps of history" pile up in total | 45 |
279
280**In words:** "each call pays for the starting prompt plus everything added
281so far, so the history term grows with the square of the number of steps."
282
283**With the numbers:** 10 × 2,000 + 500 × 45 = 20,000 + 22,500 = **42,500 input
284tokens** for a task whose final conversation is only 7,000 tokens long:
285the tenth call re-sends 6,500 tokens, and its own 500-token step brings
286the history to 7,000.
287
288**In Python:**
289
290```python
291n, h_0, t = 10, 2000, 500
292# call i re-sends h_0 and i - 1 steps
293calls = [h_0 + (i - 1) * t for i in range(1, n + 1)]
294calls[0], calls[-1]  # → (2000, 6500)
295# Σ over every call
296sum(calls)  # → 42500
297# the shortcut on the right agrees
298n * h_0 + t * n * (n - 1) // 2  # → 42500
299```
300
301![Ten calls bill 42,500 input tokens in total, six times the 7,000-token conversation they end with, because every call re-sends the history](figures/primer.agents.llm.loop_tokens.svg)
302
303**Reading it:** the bars are the input tokens billed on each call. They grow
304by the same amount every step, because each call re-sends the whole history.
305The line is the running total, and it curves upwards (quadratic growth). The
306dashed line is what you might naively expect: the size of the final
307conversation, billed once. The gap between the dashed line and the curve is
308why prompt caching, trimming tool outputs and keeping loops short are the
309main cost levers (`primer.agents.cost`, `primer.agents.context`).
310
311**In code:** `loop_input_tokens` lists the tokens billed on each call of
312such a loop. `estimate_tokens` is the four-characters-per-token rule of
313thumb, and `conversation_chars` measures everything a call resends, which is
314how `ScriptedLLM` gives its fake replies realistic `Usage`.
315
316## In 20 seconds
317
318- A model call sends the full list of messages and returns one message; the
319  model keeps no state between calls.
320- A tool call is a structured *request* (`tool_use`). Your code validates it,
321  runs it, and returns a `tool_result` with the matching id. The model
322  executes nothing.
323- `stop_reason` tells you what to do next: `tool_use` means run tools and
324  call again, `end_turn` means done.
325- Every call re-sends the whole history, so input cost grows quadratically
326  with the number of steps in a loop.
327
328## Self-test questions
329
330**What exactly happens when a model "uses a tool"?**
331You send tool definitions (name, description, JSON Schema) with the
332messages. The model replies with a `tool_use` block and `stop_reason:
333tool_use`. Your code validates the arguments, runs the function, and sends a
334new request with the history plus a `tool_result` block carrying the same
335id. The model then answers or asks for another tool.
336
337**Why is the model never a security boundary for tool use?**
338It only produces requests. Everything that actually happens goes through
339your code, so validation, permissions, approvals and rate limits all belong
340there, and a manipulated model can only ever ask.
341
342**The model asks for three tools in one reply. How do you send the results?**
343Run them (concurrently if independent) and return all three `tool_result`
344blocks in a single user message, each with its own `tool_use_id`. For one
345that failed, return a `tool_result` with `is_error: true` and an actionable
346message rather than dropping it.
347
348**A 20-step agent loop costs far more than 20 times a single call. Why?**
349Each call re-sends the full, growing history, so the input billed is a sum
350that grows with the square of the number of steps. Caching the stable prefix
351and trimming tool outputs attack exactly this.
352
353**Why test agents against a scripted model at all?**
354Real models are non-deterministic and rarely misbehave on cue. A scripted
355model reproduces loops, bad arguments and injected instructions exactly, so
356the code that must handle them can be tested every time.
357
358## The papers behind this lesson
359
360- **Schick et al., *Toolformer: Language Models Can Teach Themselves to Use
361  Tools* (2023)**: https://arxiv.org/abs/2302.04761. Showed a model can learn
362  when to call external tools and how to use their results.
363  [Annotated companion](../../papers/toolformer.html)
364- **Yao et al., *ReAct: Synergizing Reasoning and Acting in Language Models*
365  (2022)**: https://arxiv.org/abs/2210.03629. The pattern of interleaving
366  reasoning with tool calls that modern agent loops descend from.
367  [Annotated companion](../../papers/react.html)
368
369## Further reading
370
371- Tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
372- Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
373- Messages API reference: https://docs.claude.com/en/api/messages
374- Python SDK: https://github.com/anthropics/anthropic-sdk-python
375- Anthropic, *Building effective agents*: https://www.anthropic.com/engineering/building-effective-agents
376"""
377
378from __future__ import annotations
379
380import itertools
381import json
382import math
383from dataclasses import dataclass, field
384from typing import Any, Callable, Literal, Protocol
385
386StopReason = Literal["end_turn", "tool_use", "max_tokens", "refusal"]
387
388# Default model for the real adapter. Opus 5 is the current default Claude
389# model for new code; pass `model=` to use another (e.g. "claude-haiku-4-5"
390# for cheap routing/classification steps; see primer.agents.cost).
391DEFAULT_MODEL = "claude-opus-5"
392
393
394def estimate_tokens(text: str) -> int:
395    """Rule-of-thumb token count: ~4 characters per token for English prose.
396
397    Good enough for budgeting demos. For real numbers, use the provider's
398    token-counting endpoint, since vocabularies differ between models.
399    https://docs.claude.com/en/docs/build-with-claude/token-counting
400    """
401    return max(1, math.ceil(len(text) / 4))
402
403
404@dataclass
405class ToolCall:
406    """One `tool_use` block: the model *asking* to run a tool."""
407
408    id: str
409    name: str
410    input: dict[str, Any]
411
412
413@dataclass
414class Usage:
415    input_tokens: int = 0
416    output_tokens: int = 0
417
418    def __iadd__(self, other: "Usage") -> "Usage":
419        self.input_tokens += other.input_tokens
420        self.output_tokens += other.output_tokens
421        return self
422
423
424@dataclass
425class LLMResponse:
426    """What an `LLM.complete` call returns, normalized across implementations.
427
428    `assistant_content` is the exact list of content blocks to append to the
429    conversation as the assistant turn. Always append *that* (not just the
430    text): with the real API it can include thinking blocks that must be sent
431    back unchanged.
432    """
433
434    text: str
435    tool_calls: list[ToolCall]
436    stop_reason: StopReason
437    usage: Usage
438    model: str
439    assistant_content: list[dict[str, Any]] = field(default_factory=list)
440
441
442class LLM(Protocol):
443    """The one method agents need."""
444
445    def complete(
446        self,
447        *,
448        system: str,
449        messages: list[dict[str, Any]],
450        tools: list[dict[str, Any]] | None = None,
451        max_tokens: int = 4096,
452    ) -> LLMResponse: ...
453
454
455# ----------------------------------------------------------------------------
456# Helpers for reading a conversation (useful inside ScriptedLLM policies)
457# ----------------------------------------------------------------------------
458
459
460def content_blocks(message: dict[str, Any]) -> list[dict[str, Any]]:
461    """A message's content as a list of blocks (a bare string becomes one text block)."""
462    c = message["content"]
463    return [{"type": "text", "text": c}] if isinstance(c, str) else list(c)
464
465
466def last_user_text(messages: list[dict[str, Any]]) -> str:
467    """Text of the most recent user message that contains text (not just tool results)."""
468    for m in reversed(messages):
469        if m["role"] != "user":
470            continue
471        texts = [b["text"] for b in content_blocks(m) if b.get("type") == "text"]
472        if texts:
473            return "\n".join(texts)
474    return ""
475
476
477def tool_results(messages: list[dict[str, Any]]) -> list[dict[str, Any]]:
478    """All `tool_result` blocks in the conversation, oldest first."""
479    return [b for m in messages if m["role"] == "user" for b in content_blocks(m) if b.get("type") == "tool_result"]
480
481
482def tool_calls_so_far(messages: list[dict[str, Any]]) -> list[dict[str, Any]]:
483    """All `tool_use` blocks the assistant has emitted, oldest first."""
484    return [b for m in messages if m["role"] == "assistant" for b in content_blocks(m) if b.get("type") == "tool_use"]
485
486
487def conversation_chars(system: str, messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None) -> int:
488    """Rough size of everything sent to the model (for token estimates)."""
489    return len(system) + len(json.dumps(messages, default=str)) + len(json.dumps(tools or []))
490
491
492# ----------------------------------------------------------------------------
493# ScriptedLLM: deterministic offline "model"
494# ----------------------------------------------------------------------------
495
496# A policy decides the next assistant turn. It returns either:
497#   * a string                       -> final text answer (stop_reason end_turn)
498#   * a ToolCall or list[ToolCall]   -> tool request(s)   (stop_reason tool_use)
499#   * (text, [ToolCall, ...])        -> text plus tool requests
500#   * an LLMResponse                 -> used as-is (e.g. to simulate a refusal)
501PolicyResult = Any
502Policy = Callable[[str, list[dict[str, Any]], list[dict[str, Any]] | None], PolicyResult]
503
504
505class ScriptedLLM:
506    """A fake LLM driven by a Python function, for offline demos and tests.
507
508    Examples:
509        >>> def policy(system, messages, tools):
510        ...     if not tool_results(messages):
511        ...         return ToolCall("t1", "get_time", {})
512        ...     return "It is " + tool_results(messages)[-1]["content"]
513        >>> llm = ScriptedLLM(policy)
514        >>> r = llm.complete(system="", messages=[{"role": "user", "content": "time?"}])
515        >>> r.stop_reason, r.tool_calls[0].name
516        ('tool_use', 'get_time')
517
518    Token usage is *estimated* from characters so cost demos have realistic
519    shapes: input grows with the whole conversation each call, which is the
520    single most important cost fact about agent loops.
521    """
522
523    def __init__(self, policy: Policy | list[PolicyResult], model: str = "scripted-large"):
524        if isinstance(policy, list):
525            # A fixed script: return each item in turn, repeating the last one.
526            script = list(policy)
527            counter = itertools.count()
528
529            def policy_fn(system, messages, tools):  # noqa: ARG001
530                return script[min(next(counter), len(script) - 1)]
531
532            self.policy: Policy = policy_fn
533        else:
534            self.policy = policy
535        self.model = model
536        self.calls = 0
537        self.total_usage = Usage()
538        self._ids = itertools.count(1)
539
540    def complete(self, *, system, messages, tools=None, max_tokens=4096) -> LLMResponse:  # noqa: ARG002
541        self.calls += 1
542        out = self.policy(system, messages, tools)
543        if isinstance(out, LLMResponse):
544            self.total_usage += out.usage
545            return out
546
547        text, calls = "", []
548        if isinstance(out, str):
549            text = out
550        elif isinstance(out, ToolCall):
551            calls = [out]
552        elif isinstance(out, tuple):
553            text, calls = out
554        elif isinstance(out, list):
555            calls = out
556        else:
557            raise TypeError(f"policy returned unsupported {type(out).__name__}")
558
559        # Give tool calls unique ids if the policy didn't bother.
560        calls = [c if c.id else ToolCall(f"toolu_{next(self._ids)}", c.name, c.input) for c in calls]
561
562        blocks: list[dict[str, Any]] = []
563        if text:
564            blocks.append({"type": "text", "text": text})
565        blocks += [{"type": "tool_use", "id": c.id, "name": c.name, "input": c.input} for c in calls]
566
567        usage = Usage(
568            input_tokens=estimate_tokens("x" * conversation_chars(system, messages, tools)),
569            output_tokens=estimate_tokens(json.dumps(blocks)),
570        )
571        self.total_usage += usage
572        return LLMResponse(
573            text=text,
574            tool_calls=calls,
575            stop_reason="tool_use" if calls else "end_turn",
576            usage=usage,
577            model=self.model,
578            assistant_content=blocks,
579        )
580
581
582# ----------------------------------------------------------------------------
583# ClaudeLLM: the real model via the Anthropic SDK
584# ----------------------------------------------------------------------------
585
586
587class ClaudeLLM:
588    """Adapter from the `LLM` interface to the Anthropic Messages API.
589
590    Requires `pip install anthropic` and credentials (`ANTHROPIC_API_KEY`, or
591    `ant auth login`). Nothing else in this repo needs it.
592
593    Args:
594        model: model id, default `claude-opus-5`.
595        fallbacks: when True (default), opts into server-side refusal
596            fallbacks, so if a safety classifier declines the request the API
597            retries on a fallback model instead of returning
598            `stop_reason == "refusal"`. Set False if your account or proxy
599            rejects the beta header. We still check for `"refusal"` either way.
600
601    Notes:
602        * Adaptive thinking is on by default for Opus 5. The response may
603          contain `thinking` blocks; `assistant_content` preserves them so the
604          agent loop can send them back unchanged, as the API requires.
605        * Tool inputs are parsed JSON dicts, never raw strings. Don't
606          string-match serialized input.
607    """
608
609    def __init__(self, model: str = DEFAULT_MODEL, fallbacks: bool = True, effort: str | None = None):
610        import anthropic  # imported lazily so the rest of the repo has no hard dependency
611
612        self._anthropic = anthropic
613        self.client = anthropic.Anthropic()
614        self.model = model
615        self.fallbacks = fallbacks
616        self.effort = effort
617
618    def complete(self, *, system, messages, tools=None, max_tokens=16000) -> LLMResponse:
619        kwargs = claude_request(
620            model=self.model, system=system, messages=messages, tools=tools,
621            max_tokens=max_tokens, effort=self.effort, fallbacks=self.fallbacks,
622        )
623        resp = self.client.messages.create(**kwargs)
624
625        blocks = [b.model_dump(exclude_none=True) for b in resp.content]
626        text = "".join(b.text for b in resp.content if b.type == "text")
627        calls = [ToolCall(b.id, b.name, dict(b.input)) for b in resp.content if b.type == "tool_use"]
628        stop = resp.stop_reason if resp.stop_reason in ("end_turn", "tool_use", "max_tokens", "refusal") else "end_turn"
629        return LLMResponse(
630            text=text,
631            tool_calls=calls,
632            stop_reason=stop,  # type: ignore[arg-type]
633            usage=Usage(resp.usage.input_tokens, resp.usage.output_tokens),
634            model=resp.model,
635            assistant_content=blocks,
636        )
637
638
639def claude_request(*, model, system, messages, tools, max_tokens, effort, fallbacks) -> dict[str, Any]:
640    """The keyword arguments for `client.messages.create`, built without touching the network.
641
642    `effort` ("low" to "max") trades thinking depth for speed and cost. Low suits
643    short, grounded tasks such as a two-sentence explanation, where default-depth
644    thinking could spend a small `max_tokens` cap before writing any text.
645    """
646    kwargs: dict[str, Any] = dict(model=model, max_tokens=max_tokens, system=system, messages=messages)
647    if tools:
648        kwargs["tools"] = tools
649    extra_body: dict[str, Any] = {}
650    if fallbacks:
651        # Server-side refusal fallback (beta): route by refusal category.
652        kwargs["extra_headers"] = {"anthropic-beta": "server-side-fallback-2026-07-01"}
653        extra_body["fallbacks"] = "default"
654    if effort:
655        extra_body["output_config"] = {"effort": effort}
656    if extra_body:
657        kwargs["extra_body"] = extra_body
658    return kwargs
659
660
661# ----------------------------------------------------------------------------
662# OllamaLLM: an open model running on your own machine
663# ----------------------------------------------------------------------------
664
665OLLAMA_HOST = "http://127.0.0.1:11434"
666
667
668def _plain_text(content: Any) -> str:
669    """Ollama messages carry plain strings; join the text of any content blocks."""
670    if isinstance(content, str):
671        return content
672    return "\n".join(b.get("text", "") or str(b.get("content", "")) for b in content if isinstance(b, dict))
673
674
675def ollama_request(*, model: str, system: str, messages: list[dict[str, Any]], max_tokens: int,
676                   temperature: float = 0.2, keep_alive: str = "30m") -> dict[str, Any]:
677    """The JSON body for Ollama's /api/chat, built without touching the network.
678
679    `think: False` matters: some small open models "think out loud" by default and
680    would spend the whole token budget reasoning instead of answering.
681    `keep_alive` keeps the weights in memory between calls, so only the first
682    call pays to load them.
683    """
684    msgs = ([{"role": "system", "content": system}] if system else []) + [
685        {"role": m["role"], "content": _plain_text(m["content"])} for m in messages
686    ]
687    return {
688        "model": model,
689        "messages": msgs,
690        "stream": False,
691        "think": False,
692        "keep_alive": keep_alive,
693        "options": {"temperature": temperature, "num_predict": max_tokens},
694    }
695
696
697def parse_ollama_reply(data: dict[str, Any]) -> LLMResponse:
698    """Normalize Ollama's /api/chat reply into an `LLMResponse`."""
699    text = (data.get("message") or {}).get("content", "").strip()
700    stop: StopReason = "max_tokens" if data.get("done_reason") == "length" else "end_turn"
701    return LLMResponse(
702        text=text,
703        tool_calls=[],
704        stop_reason=stop,
705        usage=Usage(data.get("prompt_eval_count", 0), data.get("eval_count", 0)),
706        model=data.get("model", ""),
707        assistant_content=[{"type": "text", "text": text}] if text else [],
708    )
709
710
711class OllamaLLM:
712    """Adapter from the `LLM` interface to a local model served by Ollama.
713
714    Free, private and offline: nothing leaves your machine. Text only (no tool
715    calling here). Start Ollama and download the model first
716    (`ollama pull qwen3:1.7b`).
717    """
718
719    def __init__(self, model: str = "qwen3:1.7b", host: str = OLLAMA_HOST, temperature: float = 0.2):
720        self.model, self.host, self.temperature = model, host, temperature
721
722    def complete(self, *, system, messages, tools=None, max_tokens=300) -> LLMResponse:
723        import urllib.request
724
725        if tools:
726            raise NotImplementedError("OllamaLLM here handles text only; use ClaudeLLM for tool calling")
727        body = ollama_request(model=self.model, system=system, messages=messages,
728                              max_tokens=max_tokens, temperature=self.temperature)
729        req = urllib.request.Request(f"{self.host}/api/chat", json.dumps(body).encode(), {"Content-Type": "application/json"})
730        with urllib.request.urlopen(req, timeout=120) as r:
731            return parse_ollama_reply(json.load(r))
732
733
734def tool_result_block(tool_use_id: str, content: str, is_error: bool = False) -> dict[str, Any]:
735    """Build a `tool_result` block. On failure, set `is_error=True` with an
736    actionable message the model can use to fix its next call."""
737    block: dict[str, Any] = {"type": "tool_result", "tool_use_id": tool_use_id, "content": content}
738    if is_error:
739        block["is_error"] = True
740    return block
741
742
743# ----------------------------------------------------------------------------
744# The worked example, as code
745# ----------------------------------------------------------------------------
746
747WEATHER_TOOL = {
748    "name": "get_weather",
749    "description": "Current weather for a city. Use for any question about today's weather.",
750    "input_schema": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]},
751}
752
753
754def _get_weather(city: str) -> str:
755    return {"paris": "18°C, sunny"}.get(city.lower(), "no data")
756
757
758def trace_tool_call_round_trip() -> list[dict[str, Any]]:
759    """Run the lesson's worked example and return the full transcript (4 messages)."""
760
761    def policy(system, messages, tools):
762        results = tool_results(messages)
763        if not results:  # first call: ask for the tool
764            return ToolCall("", "get_weather", {"city": "Paris"})
765        return f"It's {results[-1]['content'].replace(', ', ' and ')} in Paris."
766
767    llm = ScriptedLLM(policy)
768    messages: list[dict[str, Any]] = [{"role": "user", "content": "What's the weather in Paris?"}]
769    while True:
770        reply = llm.complete(system="You are helpful.", messages=messages, tools=[WEATHER_TOOL])
771        messages.append({"role": "assistant", "content": reply.assistant_content})
772        if reply.stop_reason != "tool_use":
773            return messages
774        # Your code, not the model, runs the tool.
775        messages.append(
776            {"role": "user", "content": [tool_result_block(c.id, _get_weather(**c.input)) for c in reply.tool_calls]}
777        )
778
779
780def loop_input_tokens(n_calls: int, first: int, per_step: int) -> list[int]:
781    """Input tokens billed on each call of a loop that re-sends its growing history."""
782    return [first + i * per_step for i in range(n_calls)]
783
784
785def figures() -> dict:
786    import matplotlib
787
788    matplotlib.use("Agg")
789    import matplotlib.pyplot as plt
790    import numpy as np
791
792    per_call = loop_input_tokens(10, 2000, 500)
793    fig, ax = plt.subplots(figsize=(6.4, 3.8))
794    x = np.arange(1, 11)
795    ax.bar(x, per_call, color="#93c5fd", label="input tokens billed on this call")
796    ax.plot(x, np.cumsum(per_call), "o-", color="#2563eb", label="running total")
797    ax.axhline(per_call[-1] + 500, ls="--", color="#9ca3af", label="final conversation size (billed once?)")
798    ax.set_xlabel("model call in the loop")
799    ax.set_ylabel("input tokens")
800    ax.set_title("Every call re-sends the whole history")
801    ax.legend(frameon=False)
802    return {"loop_tokens": fig}
803
804
805def demo() -> None:
806    from primer._show import banner, say, table, takeaway
807
808    banner("1. One tool call, traced message by message")
809    for i, m in enumerate(trace_tool_call_round_trip(), 1):
810        for b in content_blocks(m):
811            detail = b.get("text") or (f"{b['name']}({b['input']}) id={b['id']}" if b["type"] == "tool_use" else f"result for {b['tool_use_id']}: {b['content']}")
812            print(f"  {i}. {m['role']:9s} {b['type']:11s} {detail}")
813    print()
814    takeaway("The model asked; your code ran the tool; the model answered. The model executed nothing.")
815
816    banner("2. Every call pays for the whole conversation")
817    per_call = loop_input_tokens(10, 2000, 500)
818    table(["call", "input tokens", "running total"], [(i + 1, t, sum(per_call[: i + 1])) for i, t in enumerate(per_call)])
819    say("10 calls bill 42,500 input tokens for a conversation that ends 7,000 tokens long.")
820
821
822if __name__ == "__main__":
823    demo()
Level 3: the code, function by function.
StopReason = typing.Literal['end_turn', 'tool_use', 'max_tokens', 'refusal']
DEFAULT_MODEL = 'claude-opus-5'
def estimate_tokens(text: str) -> int: on GitHub
395def estimate_tokens(text: str) -> int:
396    """Rule-of-thumb token count: ~4 characters per token for English prose.
397
398    Good enough for budgeting demos. For real numbers, use the provider's
399    token-counting endpoint, since vocabularies differ between models.
400    https://docs.claude.com/en/docs/build-with-claude/token-counting
401    """
402    return max(1, math.ceil(len(text) / 4))

Rule-of-thumb token count: ~4 characters per token for English prose.

Good enough for budgeting demos. For real numbers, use the provider's token-counting endpoint, since vocabularies differ between models. https://docs.claude.com/en/docs/build-with-claude/token-counting

@dataclass
class ToolCall: on GitHub
405@dataclass
406class ToolCall:
407    """One `tool_use` block: the model *asking* to run a tool."""
408
409    id: str
410    name: str
411    input: dict[str, Any]

One tool_use block: the model asking to run a tool.

ToolCall(id: str, name: str, input: dict[str, typing.Any])
id: str
name: str
input: dict[str, typing.Any]
@dataclass
class Usage: on GitHub
414@dataclass
415class Usage:
416    input_tokens: int = 0
417    output_tokens: int = 0
418
419    def __iadd__(self, other: "Usage") -> "Usage":
420        self.input_tokens += other.input_tokens
421        self.output_tokens += other.output_tokens
422        return self
Usage(input_tokens: int = 0, output_tokens: int = 0)
input_tokens: int = 0
output_tokens: int = 0
@dataclass
class LLMResponse: on GitHub
425@dataclass
426class LLMResponse:
427    """What an `LLM.complete` call returns, normalized across implementations.
428
429    `assistant_content` is the exact list of content blocks to append to the
430    conversation as the assistant turn. Always append *that* (not just the
431    text): with the real API it can include thinking blocks that must be sent
432    back unchanged.
433    """
434
435    text: str
436    tool_calls: list[ToolCall]
437    stop_reason: StopReason
438    usage: Usage
439    model: str
440    assistant_content: list[dict[str, Any]] = field(default_factory=list)

What an LLM.complete call returns, normalized across implementations.

assistant_content is the exact list of content blocks to append to the conversation as the assistant turn. Always append that (not just the text): with the real API it can include thinking blocks that must be sent back unchanged.

LLMResponse( text: str, tool_calls: list[ToolCall], stop_reason: Literal['end_turn', 'tool_use', 'max_tokens', 'refusal'], usage: Usage, model: str, assistant_content: list[dict[str, typing.Any]] = <factory>)
text: str
stop_reason: Literal['end_turn', 'tool_use', 'max_tokens', 'refusal']
usage: Usage
model: str
assistant_content: list[dict[str, typing.Any]]
class LLM(typing.Protocol): on GitHub
443class LLM(Protocol):
444    """The one method agents need."""
445
446    def complete(
447        self,
448        *,
449        system: str,
450        messages: list[dict[str, Any]],
451        tools: list[dict[str, Any]] | None = None,
452        max_tokens: int = 4096,
453    ) -> LLMResponse: ...

The one method agents need.

LLM(*args, **kwargs)
1771def _no_init_or_replace_init(self, *args, **kwargs):
1772    cls = type(self)
1773
1774    if cls._is_protocol:
1775        raise TypeError('Protocols cannot be instantiated')
1776
1777    # Already using a custom `__init__`. No need to calculate correct
1778    # `__init__` to call. This can lead to RecursionError. See bpo-45121.
1779    if cls.__init__ is not _no_init_or_replace_init:
1780        return
1781
1782    # Initially, `__init__` of a protocol subclass is set to `_no_init_or_replace_init`.
1783    # The first instantiation of the subclass will call `_no_init_or_replace_init` which
1784    # searches for a proper new `__init__` in the MRO. The new `__init__`
1785    # replaces the subclass' old `__init__` (ie `_no_init_or_replace_init`). Subsequent
1786    # instantiation of the protocol subclass will thus use the new
1787    # `__init__` and no longer call `_no_init_or_replace_init`.
1788    for base in cls.__mro__:
1789        init = base.__dict__.get('__init__', _no_init_or_replace_init)
1790        if init is not _no_init_or_replace_init:
1791            cls.__init__ = init
1792            break
1793    else:
1794        # should not happen
1795        cls.__init__ = object.__init__
1796
1797    cls.__init__(self, *args, **kwargs)
def complete( self, *, system: str, messages: list[dict[str, typing.Any]], tools: list[dict[str, typing.Any]] | None = None, max_tokens: int = 4096) -> LLMResponse: on GitHub
446    def complete(
447        self,
448        *,
449        system: str,
450        messages: list[dict[str, Any]],
451        tools: list[dict[str, Any]] | None = None,
452        max_tokens: int = 4096,
453    ) -> LLMResponse: ...
def content_blocks(message: dict[str, typing.Any]) -> list[dict[str, typing.Any]]: on GitHub
461def content_blocks(message: dict[str, Any]) -> list[dict[str, Any]]:
462    """A message's content as a list of blocks (a bare string becomes one text block)."""
463    c = message["content"]
464    return [{"type": "text", "text": c}] if isinstance(c, str) else list(c)

A message's content as a list of blocks (a bare string becomes one text block).

def last_user_text(messages: list[dict[str, typing.Any]]) -> str: on GitHub
467def last_user_text(messages: list[dict[str, Any]]) -> str:
468    """Text of the most recent user message that contains text (not just tool results)."""
469    for m in reversed(messages):
470        if m["role"] != "user":
471            continue
472        texts = [b["text"] for b in content_blocks(m) if b.get("type") == "text"]
473        if texts:
474            return "\n".join(texts)
475    return ""

Text of the most recent user message that contains text (not just tool results).

def tool_results(messages: list[dict[str, typing.Any]]) -> list[dict[str, typing.Any]]: on GitHub
478def tool_results(messages: list[dict[str, Any]]) -> list[dict[str, Any]]:
479    """All `tool_result` blocks in the conversation, oldest first."""
480    return [b for m in messages if m["role"] == "user" for b in content_blocks(m) if b.get("type") == "tool_result"]

All tool_result blocks in the conversation, oldest first.

def tool_calls_so_far(messages: list[dict[str, typing.Any]]) -> list[dict[str, typing.Any]]: on GitHub
483def tool_calls_so_far(messages: list[dict[str, Any]]) -> list[dict[str, Any]]:
484    """All `tool_use` blocks the assistant has emitted, oldest first."""
485    return [b for m in messages if m["role"] == "assistant" for b in content_blocks(m) if b.get("type") == "tool_use"]

All tool_use blocks the assistant has emitted, oldest first.

def conversation_chars( system: str, messages: list[dict[str, typing.Any]], tools: list[dict[str, typing.Any]] | None) -> int: on GitHub
488def conversation_chars(system: str, messages: list[dict[str, Any]], tools: list[dict[str, Any]] | None) -> int:
489    """Rough size of everything sent to the model (for token estimates)."""
490    return len(system) + len(json.dumps(messages, default=str)) + len(json.dumps(tools or []))

Rough size of everything sent to the model (for token estimates).

PolicyResult = typing.Any
Policy = typing.Callable[[str, list[dict[str, typing.Any]], list[dict[str, typing.Any]] | None], typing.Any]
class ScriptedLLM: on GitHub
506class ScriptedLLM:
507    """A fake LLM driven by a Python function, for offline demos and tests.
508
509    Examples:
510        >>> def policy(system, messages, tools):
511        ...     if not tool_results(messages):
512        ...         return ToolCall("t1", "get_time", {})
513        ...     return "It is " + tool_results(messages)[-1]["content"]
514        >>> llm = ScriptedLLM(policy)
515        >>> r = llm.complete(system="", messages=[{"role": "user", "content": "time?"}])
516        >>> r.stop_reason, r.tool_calls[0].name
517        ('tool_use', 'get_time')
518
519    Token usage is *estimated* from characters so cost demos have realistic
520    shapes: input grows with the whole conversation each call, which is the
521    single most important cost fact about agent loops.
522    """
523
524    def __init__(self, policy: Policy | list[PolicyResult], model: str = "scripted-large"):
525        if isinstance(policy, list):
526            # A fixed script: return each item in turn, repeating the last one.
527            script = list(policy)
528            counter = itertools.count()
529
530            def policy_fn(system, messages, tools):  # noqa: ARG001
531                return script[min(next(counter), len(script) - 1)]
532
533            self.policy: Policy = policy_fn
534        else:
535            self.policy = policy
536        self.model = model
537        self.calls = 0
538        self.total_usage = Usage()
539        self._ids = itertools.count(1)
540
541    def complete(self, *, system, messages, tools=None, max_tokens=4096) -> LLMResponse:  # noqa: ARG002
542        self.calls += 1
543        out = self.policy(system, messages, tools)
544        if isinstance(out, LLMResponse):
545            self.total_usage += out.usage
546            return out
547
548        text, calls = "", []
549        if isinstance(out, str):
550            text = out
551        elif isinstance(out, ToolCall):
552            calls = [out]
553        elif isinstance(out, tuple):
554            text, calls = out
555        elif isinstance(out, list):
556            calls = out
557        else:
558            raise TypeError(f"policy returned unsupported {type(out).__name__}")
559
560        # Give tool calls unique ids if the policy didn't bother.
561        calls = [c if c.id else ToolCall(f"toolu_{next(self._ids)}", c.name, c.input) for c in calls]
562
563        blocks: list[dict[str, Any]] = []
564        if text:
565            blocks.append({"type": "text", "text": text})
566        blocks += [{"type": "tool_use", "id": c.id, "name": c.name, "input": c.input} for c in calls]
567
568        usage = Usage(
569            input_tokens=estimate_tokens("x" * conversation_chars(system, messages, tools)),
570            output_tokens=estimate_tokens(json.dumps(blocks)),
571        )
572        self.total_usage += usage
573        return LLMResponse(
574            text=text,
575            tool_calls=calls,
576            stop_reason="tool_use" if calls else "end_turn",
577            usage=usage,
578            model=self.model,
579            assistant_content=blocks,
580        )

A fake LLM driven by a Python function, for offline demos and tests.

Examples:

>>> def policy(system, messages, tools):
...     if not tool_results(messages):
...         return ToolCall("t1", "get_time", {})
...     return "It is " + tool_results(messages)[-1]["content"]
>>> llm = ScriptedLLM(policy)
>>> r = llm.complete(system="", messages=[{"role": "user", "content": "time?"}])
>>> r.stop_reason, r.tool_calls[0].name
('tool_use', 'get_time')

Token usage is estimated from characters so cost demos have realistic shapes: input grows with the whole conversation each call, which is the single most important cost fact about agent loops.

ScriptedLLM( policy: Union[Callable[[str, list[dict[str, Any]], list[dict[str, Any]] | None], Any], list[Any]], model: str = 'scripted-large') on GitHub
524    def __init__(self, policy: Policy | list[PolicyResult], model: str = "scripted-large"):
525        if isinstance(policy, list):
526            # A fixed script: return each item in turn, repeating the last one.
527            script = list(policy)
528            counter = itertools.count()
529
530            def policy_fn(system, messages, tools):  # noqa: ARG001
531                return script[min(next(counter), len(script) - 1)]
532
533            self.policy: Policy = policy_fn
534        else:
535            self.policy = policy
536        self.model = model
537        self.calls = 0
538        self.total_usage = Usage()
539        self._ids = itertools.count(1)
model
calls
def complete( self, *, system, messages, tools=None, max_tokens=4096) -> LLMResponse: on GitHub
541    def complete(self, *, system, messages, tools=None, max_tokens=4096) -> LLMResponse:  # noqa: ARG002
542        self.calls += 1
543        out = self.policy(system, messages, tools)
544        if isinstance(out, LLMResponse):
545            self.total_usage += out.usage
546            return out
547
548        text, calls = "", []
549        if isinstance(out, str):
550            text = out
551        elif isinstance(out, ToolCall):
552            calls = [out]
553        elif isinstance(out, tuple):
554            text, calls = out
555        elif isinstance(out, list):
556            calls = out
557        else:
558            raise TypeError(f"policy returned unsupported {type(out).__name__}")
559
560        # Give tool calls unique ids if the policy didn't bother.
561        calls = [c if c.id else ToolCall(f"toolu_{next(self._ids)}", c.name, c.input) for c in calls]
562
563        blocks: list[dict[str, Any]] = []
564        if text:
565            blocks.append({"type": "text", "text": text})
566        blocks += [{"type": "tool_use", "id": c.id, "name": c.name, "input": c.input} for c in calls]
567
568        usage = Usage(
569            input_tokens=estimate_tokens("x" * conversation_chars(system, messages, tools)),
570            output_tokens=estimate_tokens(json.dumps(blocks)),
571        )
572        self.total_usage += usage
573        return LLMResponse(
574            text=text,
575            tool_calls=calls,
576            stop_reason="tool_use" if calls else "end_turn",
577            usage=usage,
578            model=self.model,
579            assistant_content=blocks,
580        )
class ClaudeLLM: on GitHub
588class ClaudeLLM:
589    """Adapter from the `LLM` interface to the Anthropic Messages API.
590
591    Requires `pip install anthropic` and credentials (`ANTHROPIC_API_KEY`, or
592    `ant auth login`). Nothing else in this repo needs it.
593
594    Args:
595        model: model id, default `claude-opus-5`.
596        fallbacks: when True (default), opts into server-side refusal
597            fallbacks, so if a safety classifier declines the request the API
598            retries on a fallback model instead of returning
599            `stop_reason == "refusal"`. Set False if your account or proxy
600            rejects the beta header. We still check for `"refusal"` either way.
601
602    Notes:
603        * Adaptive thinking is on by default for Opus 5. The response may
604          contain `thinking` blocks; `assistant_content` preserves them so the
605          agent loop can send them back unchanged, as the API requires.
606        * Tool inputs are parsed JSON dicts, never raw strings. Don't
607          string-match serialized input.
608    """
609
610    def __init__(self, model: str = DEFAULT_MODEL, fallbacks: bool = True, effort: str | None = None):
611        import anthropic  # imported lazily so the rest of the repo has no hard dependency
612
613        self._anthropic = anthropic
614        self.client = anthropic.Anthropic()
615        self.model = model
616        self.fallbacks = fallbacks
617        self.effort = effort
618
619    def complete(self, *, system, messages, tools=None, max_tokens=16000) -> LLMResponse:
620        kwargs = claude_request(
621            model=self.model, system=system, messages=messages, tools=tools,
622            max_tokens=max_tokens, effort=self.effort, fallbacks=self.fallbacks,
623        )
624        resp = self.client.messages.create(**kwargs)
625
626        blocks = [b.model_dump(exclude_none=True) for b in resp.content]
627        text = "".join(b.text for b in resp.content if b.type == "text")
628        calls = [ToolCall(b.id, b.name, dict(b.input)) for b in resp.content if b.type == "tool_use"]
629        stop = resp.stop_reason if resp.stop_reason in ("end_turn", "tool_use", "max_tokens", "refusal") else "end_turn"
630        return LLMResponse(
631            text=text,
632            tool_calls=calls,
633            stop_reason=stop,  # type: ignore[arg-type]
634            usage=Usage(resp.usage.input_tokens, resp.usage.output_tokens),
635            model=resp.model,
636            assistant_content=blocks,
637        )

Adapter from the LLM interface to the Anthropic Messages API.

Requires pip install anthropic and credentials (ANTHROPIC_API_KEY, or ant auth login). Nothing else in this repo needs it.

Arguments:

  • model: model id, default claude-opus-5.
  • fallbacks: when True (default), opts into server-side refusal fallbacks, so if a safety classifier declines the request the API retries on a fallback model instead of returning stop_reason == "refusal". Set False if your account or proxy rejects the beta header. We still check for "refusal" either way.

Notes:

  • Adaptive thinking is on by default for Opus 5. The response may contain thinking blocks; assistant_content preserves them so the agent loop can send them back unchanged, as the API requires.
  • Tool inputs are parsed JSON dicts, never raw strings. Don't string-match serialized input.
ClaudeLLM( model: str = 'claude-opus-5', fallbacks: bool = True, effort: str | None = None) on GitHub
610    def __init__(self, model: str = DEFAULT_MODEL, fallbacks: bool = True, effort: str | None = None):
611        import anthropic  # imported lazily so the rest of the repo has no hard dependency
612
613        self._anthropic = anthropic
614        self.client = anthropic.Anthropic()
615        self.model = model
616        self.fallbacks = fallbacks
617        self.effort = effort
client
model
fallbacks
effort
def complete( self, *, system, messages, tools=None, max_tokens=16000) -> LLMResponse: on GitHub
619    def complete(self, *, system, messages, tools=None, max_tokens=16000) -> LLMResponse:
620        kwargs = claude_request(
621            model=self.model, system=system, messages=messages, tools=tools,
622            max_tokens=max_tokens, effort=self.effort, fallbacks=self.fallbacks,
623        )
624        resp = self.client.messages.create(**kwargs)
625
626        blocks = [b.model_dump(exclude_none=True) for b in resp.content]
627        text = "".join(b.text for b in resp.content if b.type == "text")
628        calls = [ToolCall(b.id, b.name, dict(b.input)) for b in resp.content if b.type == "tool_use"]
629        stop = resp.stop_reason if resp.stop_reason in ("end_turn", "tool_use", "max_tokens", "refusal") else "end_turn"
630        return LLMResponse(
631            text=text,
632            tool_calls=calls,
633            stop_reason=stop,  # type: ignore[arg-type]
634            usage=Usage(resp.usage.input_tokens, resp.usage.output_tokens),
635            model=resp.model,
636            assistant_content=blocks,
637        )
def claude_request( *, model, system, messages, tools, max_tokens, effort, fallbacks) -> dict[str, typing.Any]: on GitHub
640def claude_request(*, model, system, messages, tools, max_tokens, effort, fallbacks) -> dict[str, Any]:
641    """The keyword arguments for `client.messages.create`, built without touching the network.
642
643    `effort` ("low" to "max") trades thinking depth for speed and cost. Low suits
644    short, grounded tasks such as a two-sentence explanation, where default-depth
645    thinking could spend a small `max_tokens` cap before writing any text.
646    """
647    kwargs: dict[str, Any] = dict(model=model, max_tokens=max_tokens, system=system, messages=messages)
648    if tools:
649        kwargs["tools"] = tools
650    extra_body: dict[str, Any] = {}
651    if fallbacks:
652        # Server-side refusal fallback (beta): route by refusal category.
653        kwargs["extra_headers"] = {"anthropic-beta": "server-side-fallback-2026-07-01"}
654        extra_body["fallbacks"] = "default"
655    if effort:
656        extra_body["output_config"] = {"effort": effort}
657    if extra_body:
658        kwargs["extra_body"] = extra_body
659    return kwargs

The keyword arguments for client.messages.create, built without touching the network.

effort ("low" to "max") trades thinking depth for speed and cost. Low suits short, grounded tasks such as a two-sentence explanation, where default-depth thinking could spend a small max_tokens cap before writing any text.

OLLAMA_HOST = 'http://127.0.0.1:11434'
def ollama_request( *, model: str, system: str, messages: list[dict[str, typing.Any]], max_tokens: int, temperature: float = 0.2, keep_alive: str = '30m') -> dict[str, typing.Any]: on GitHub
676def ollama_request(*, model: str, system: str, messages: list[dict[str, Any]], max_tokens: int,
677                   temperature: float = 0.2, keep_alive: str = "30m") -> dict[str, Any]:
678    """The JSON body for Ollama's /api/chat, built without touching the network.
679
680    `think: False` matters: some small open models "think out loud" by default and
681    would spend the whole token budget reasoning instead of answering.
682    `keep_alive` keeps the weights in memory between calls, so only the first
683    call pays to load them.
684    """
685    msgs = ([{"role": "system", "content": system}] if system else []) + [
686        {"role": m["role"], "content": _plain_text(m["content"])} for m in messages
687    ]
688    return {
689        "model": model,
690        "messages": msgs,
691        "stream": False,
692        "think": False,
693        "keep_alive": keep_alive,
694        "options": {"temperature": temperature, "num_predict": max_tokens},
695    }

The JSON body for Ollama's /api/chat, built without touching the network.

think: False matters: some small open models "think out loud" by default and would spend the whole token budget reasoning instead of answering. keep_alive keeps the weights in memory between calls, so only the first call pays to load them.

def parse_ollama_reply(data: dict[str, typing.Any]) -> LLMResponse: on GitHub
698def parse_ollama_reply(data: dict[str, Any]) -> LLMResponse:
699    """Normalize Ollama's /api/chat reply into an `LLMResponse`."""
700    text = (data.get("message") or {}).get("content", "").strip()
701    stop: StopReason = "max_tokens" if data.get("done_reason") == "length" else "end_turn"
702    return LLMResponse(
703        text=text,
704        tool_calls=[],
705        stop_reason=stop,
706        usage=Usage(data.get("prompt_eval_count", 0), data.get("eval_count", 0)),
707        model=data.get("model", ""),
708        assistant_content=[{"type": "text", "text": text}] if text else [],
709    )

Normalize Ollama's /api/chat reply into an LLMResponse.

class OllamaLLM: on GitHub
712class OllamaLLM:
713    """Adapter from the `LLM` interface to a local model served by Ollama.
714
715    Free, private and offline: nothing leaves your machine. Text only (no tool
716    calling here). Start Ollama and download the model first
717    (`ollama pull qwen3:1.7b`).
718    """
719
720    def __init__(self, model: str = "qwen3:1.7b", host: str = OLLAMA_HOST, temperature: float = 0.2):
721        self.model, self.host, self.temperature = model, host, temperature
722
723    def complete(self, *, system, messages, tools=None, max_tokens=300) -> LLMResponse:
724        import urllib.request
725
726        if tools:
727            raise NotImplementedError("OllamaLLM here handles text only; use ClaudeLLM for tool calling")
728        body = ollama_request(model=self.model, system=system, messages=messages,
729                              max_tokens=max_tokens, temperature=self.temperature)
730        req = urllib.request.Request(f"{self.host}/api/chat", json.dumps(body).encode(), {"Content-Type": "application/json"})
731        with urllib.request.urlopen(req, timeout=120) as r:
732            return parse_ollama_reply(json.load(r))

Adapter from the LLM interface to a local model served by Ollama.

Free, private and offline: nothing leaves your machine. Text only (no tool calling here). Start Ollama and download the model first (ollama pull qwen3:1.7b).

OllamaLLM( model: str = 'qwen3:1.7b', host: str = 'http://127.0.0.1:11434', temperature: float = 0.2) on GitHub
720    def __init__(self, model: str = "qwen3:1.7b", host: str = OLLAMA_HOST, temperature: float = 0.2):
721        self.model, self.host, self.temperature = model, host, temperature
def complete( self, *, system, messages, tools=None, max_tokens=300) -> LLMResponse: on GitHub
723    def complete(self, *, system, messages, tools=None, max_tokens=300) -> LLMResponse:
724        import urllib.request
725
726        if tools:
727            raise NotImplementedError("OllamaLLM here handles text only; use ClaudeLLM for tool calling")
728        body = ollama_request(model=self.model, system=system, messages=messages,
729                              max_tokens=max_tokens, temperature=self.temperature)
730        req = urllib.request.Request(f"{self.host}/api/chat", json.dumps(body).encode(), {"Content-Type": "application/json"})
731        with urllib.request.urlopen(req, timeout=120) as r:
732            return parse_ollama_reply(json.load(r))
def tool_result_block( tool_use_id: str, content: str, is_error: bool = False) -> dict[str, typing.Any]: on GitHub
735def tool_result_block(tool_use_id: str, content: str, is_error: bool = False) -> dict[str, Any]:
736    """Build a `tool_result` block. On failure, set `is_error=True` with an
737    actionable message the model can use to fix its next call."""
738    block: dict[str, Any] = {"type": "tool_result", "tool_use_id": tool_use_id, "content": content}
739    if is_error:
740        block["is_error"] = True
741    return block

Build a tool_result block. On failure, set is_error=True with an actionable message the model can use to fix its next call.

WEATHER_TOOL = {'name': 'get_weather', 'description': "Current weather for a city. Use for any question about today's weather.", 'input_schema': {'type': 'object', 'properties': {'city': {'type': 'string'}}, 'required': ['city']}}
def trace_tool_call_round_trip() -> list[dict[str, typing.Any]]: on GitHub
759def trace_tool_call_round_trip() -> list[dict[str, Any]]:
760    """Run the lesson's worked example and return the full transcript (4 messages)."""
761
762    def policy(system, messages, tools):
763        results = tool_results(messages)
764        if not results:  # first call: ask for the tool
765            return ToolCall("", "get_weather", {"city": "Paris"})
766        return f"It's {results[-1]['content'].replace(', ', ' and ')} in Paris."
767
768    llm = ScriptedLLM(policy)
769    messages: list[dict[str, Any]] = [{"role": "user", "content": "What's the weather in Paris?"}]
770    while True:
771        reply = llm.complete(system="You are helpful.", messages=messages, tools=[WEATHER_TOOL])
772        messages.append({"role": "assistant", "content": reply.assistant_content})
773        if reply.stop_reason != "tool_use":
774            return messages
775        # Your code, not the model, runs the tool.
776        messages.append(
777            {"role": "user", "content": [tool_result_block(c.id, _get_weather(**c.input)) for c in reply.tool_calls]}
778        )

Run the lesson's worked example and return the full transcript (4 messages).

def loop_input_tokens(n_calls: int, first: int, per_step: int) -> list[int]: on GitHub
781def loop_input_tokens(n_calls: int, first: int, per_step: int) -> list[int]:
782    """Input tokens billed on each call of a loop that re-sends its growing history."""
783    return [first + i * per_step for i in range(n_calls)]

Input tokens billed on each call of a loop that re-sends its growing history.

def figures() -> dict: on GitHub
786def figures() -> dict:
787    import matplotlib
788
789    matplotlib.use("Agg")
790    import matplotlib.pyplot as plt
791    import numpy as np
792
793    per_call = loop_input_tokens(10, 2000, 500)
794    fig, ax = plt.subplots(figsize=(6.4, 3.8))
795    x = np.arange(1, 11)
796    ax.bar(x, per_call, color="#93c5fd", label="input tokens billed on this call")
797    ax.plot(x, np.cumsum(per_call), "o-", color="#2563eb", label="running total")
798    ax.axhline(per_call[-1] + 500, ls="--", color="#9ca3af", label="final conversation size (billed once?)")
799    ax.set_xlabel("model call in the loop")
800    ax.set_ylabel("input tokens")
801    ax.set_title("Every call re-sends the whole history")
802    ax.legend(frameon=False)
803    return {"loop_tokens": fig}
def demo() -> None: on GitHub
806def demo() -> None:
807    from primer._show import banner, say, table, takeaway
808
809    banner("1. One tool call, traced message by message")
810    for i, m in enumerate(trace_tool_call_round_trip(), 1):
811        for b in content_blocks(m):
812            detail = b.get("text") or (f"{b['name']}({b['input']}) id={b['id']}" if b["type"] == "tool_use" else f"result for {b['tool_use_id']}: {b['content']}")
813            print(f"  {i}. {m['role']:9s} {b['type']:11s} {detail}")
814    print()
815    takeaway("The model asked; your code ran the tool; the model answered. The model executed nothing.")
816
817    banner("2. Every call pays for the whole conversation")
818    per_call = loop_input_tokens(10, 2000, 500)
819    table(["call", "input tokens", "running total"], [(i + 1, t, sum(per_call[: i + 1])) for i, t in enumerate(per_call)])
820    say("10 calls bill 42,500 input tokens for a conversation that ends 7,000 tokens long.")