Toolformer, annotated
How to read this page
- Any dotted word explains itself when you hover, tab to, or tap it, and so does every symbol in every equation.
- The filter calculator is the heart of the paper: set the probabilities and watch a candidate tool call get kept or thrown away.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. Tool design in code is covered in the tools lesson.
Abstract
“We introduce Toolformer, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction.”Schick et al. (2023), Abstract. Read the original
Everyday picture
A bright student who is bad at long division and does not know today's date. Give them a calculator and a calendar, and they still have to learn when reaching for each one actually helps. Toolformer is a way for a language model to learn that by itself: try tool calls on ordinary text, and keep only the ones whose answers made the rest of the text easier to predict.
What the paper claims
- A model can learn to use a calculator, a question-answering system, a search engine, a translator and a calendar with only a handful of human-written examples per tool.
- Learning is self-supervised: the model's own prediction loss decides which tool calls were useful.
- A 6.7-billion-parameter model with tools beat the 175-billion-parameter GPT-3 on several zero-shot benchmarks, without losing its ordinary language-modelling ability.
Why it matters today
Today's models are taught tool calling mostly through large, curated training sets, but Toolformer's core test (does the tool's answer actually help?) remains a useful way to think about when an agent should call a tool at all.
1 Introduction · original
Everyday picture
Large models are paradoxical: they write essays and code, yet fumble arithmetic, forget what year it is, and invent facts. Much smaller, simpler tools do those jobs perfectly. The obvious fix is to let the model call the tools. Earlier attempts needed lots of human annotation showing when to call what, or worked only for one task decided in advance.
The two requirements the paper sets
- Learn tool use without much human annotation, partly because what a human finds useful may differ from what helps the model.
- Keep the model general: it decides for itself when and how to use which tool, on any text.
2 Approach · original
Everyday picture
Take a huge pile of ordinary web text. For each passage, let the model suggest a few places where a tool call might help and what it would ask. Actually run those calls. Then keep a call only if seeing its answer made the model better at predicting the words that follow. Finally, fine-tune the model on the passages with the kept calls written in. The model has now seen thousands of examples of useful tool calls, all chosen by its own predictions.
Hover or tap any block in the redrawn pipeline.
Hover or tap a block, left to right.
Reading it: ordinary text enters on the left. Step 1, the model itself proposes candidate positions and calls. Step 2 runs them for real. Step 3 is the judge: each call is scored by how much its result lowers the model's loss on the next few words, and only calls above a threshold survive. Step 4 fine-tunes the same model on text with the surviving calls written in. The dashed loop is the point: no human decided which calls were good, the model's own predictions did.
Writing a tool call as text
Tiny example
A tool call is written straight into the text, between special markers. The paper uses the characters “[”, “]” and “->” so the model's vocabulary needs no new tokens. An example written for this page:
The train covers 180 km in 1.5 hours, an average speed of [Calculator(180 / 1.5) -> 120.0] 120 km/h.
In words: “a call is written as the tool's name and its input between markers; the version with a result also includes an arrow followed by what the tool returned.”
With the numbers: ac = Calculator, ic = “180 / 1.5”, r = “120.0”. So e(c) is “[Calculator(180 / 1.5)]” and e(c, r) is “[Calculator(180 / 1.5) -> 120.0]”.
In Python:
# the tool and its input
a_c, i_c = "Calculator", "180 / 1.5"
# what the tool returns
r = str(180 / 1.5)
# e(c)
print(f"[{a_c}({i_c})]") # → [Calculator(180 / 1.5)]
# e(c, r)
print(f"[{a_c}({i_c}) -> {r}]") # → [Calculator(180 / 1.5) -> 120.0]
Why it matters
Writing calls as plain text means any language model can emit them, and any text corpus can be annotated with them. Modern APIs return calls as structured blocks instead, validated against a schema, but the principle is the same; see talking to a model.
Sampling candidate calls · original
Everyday picture
Show the model a few examples of text with helpful calls written in (a short prompt per tool), then give it a new passage and ask, at every position, “how likely would you be to open a tool call right here?” Positions where that probability is high become candidates, and the model writes a few possible calls at each.
In words: “for every position i, ask the model how likely it is to start a tool call there, given the prompt and the text so far; keep the positions where that is above a threshold (at most the top k), and sample up to m candidate calls at each.”
With the numbers: if τs = 0.05 and the model gives 0.31 to opening a call right after “an average speed of” but 0.002 after “The train”, only the first position is kept.
In Python:
tau_s = 0.05
# p_i after each prefix
p = {"an average speed of": 0.31, "The train": 0.002}
# I = { i | p_i > τ_s }
I = [i for i in p if p[i] > tau_s]
I # → ['an average speed of']
Filtering: keep only calls that help · original
“Intuitively, an API call is helpful to M if providing it with both the input and the output of this call makes it easier for the model to predict future tokens.”Schick et al. (2023), §2
Everyday picture
Cover the rest of a sentence and ask the student to guess the next few words three ways: with no calculator, with the calculator's question but not its answer, and with the calculator's answer. If the answer makes the guesses much better than either alternative, the calculator was worth using at that spot. Comparing against “question but no answer” is the clever part: it rules out calls that help merely by rephrasing the text.
Tiny example
Score the next three tokens after the call position, weighting nearer tokens more (weights 0.333, 0.267 and 0.2). The model's probability for each correct next token:
| Condition | token 1 | token 2 | token 3 | weighted loss |
|---|---|---|---|---|
| no call at all | 0.02 | 0.30 | 0.50 | 1.764 |
| call, but no result | 0.03 | 0.30 | 0.50 | 1.629 |
| call with its result | 0.90 | 0.60 | 0.55 | 0.291 |
For example, the first row is 0.333 × (−ln 0.02) + 0.267 × (−ln 0.30) + 0.2 × (−ln 0.50) = 0.333 × 3.912 + 0.267 × 1.204 + 0.2 × 0.693 = 1.304 + 0.321 + 0.139 = 1.764. The best alternative to the full call is the smaller of the first two rows, 1.629. The call with its result improves on that by 1.629 − 0.291 = 1.338. With a threshold of 1.0, the call is kept.
The math
In words: “measure how surprised the model is by the next few tokens, counting nearer tokens more, in three situations; keep the call only if giving the model the call and its result beats both no call and a result-less call by at least τf.”
With the numbers: the raw weights for t = 0, 1, 2, 3, 4 are 1, 0.8, 0.6, 0.4, 0.2 (then 0), summing to 3.0, so the normalised weights are 0.333, 0.267, 0.2, 0.133, 0.067. Only the next five tokens count at all. L⁻ = min(1.764, 1.629) = 1.629, L⁺ = 0.291, difference 1.338 ≥ τf = 1.0: keep.
In Python:
import math
# w̃_t = max(0, 1 − 0.2 t)
w_raw = [max(0.0, 1 - 0.2 * t) for t in range(7)]
[round(x, 1) for x in w_raw], round(sum(w_raw), 1) # → ([1.0, 0.8, 0.6, 0.4, 0.2, 0.0, 0.0], 3.0)
# w_t = w̃_t / Σ w̃_s
w = [x / sum(w_raw) for x in w_raw]
[round(x, 3) for x in w[:5]] # → [0.333, 0.267, 0.2, 0.133, 0.067]
# L_i = −Σ w_(j−i) · log p_M(x_j | ...)
def L(probs):
return -sum(w[t] * math.log(p) for t, p in enumerate(probs))
# L_i(ε): no call at all
L_none = L([0.02, 0.30, 0.50])
# L_i(e(c_i, ε)): call, but no result
L_bare = L([0.03, 0.30, 0.50])
# L_i⁺ = L_i(e(c_i, r_i)): call with its result
L_plus = L([0.90, 0.60, 0.55])
round(L_none, 3), round(L_bare, 3), round(L_plus, 3) # → (1.764, 1.629, 0.291)
# L_i⁻
L_minus = min(L_none, L_bare)
round(L_minus - L_plus, 3) # → 1.338
tau_f = 1.0
# keep the call
L_minus - L_plus >= tau_f # → True
Try it: the filter
Reading it: each slider is the model's probability for one of the next three correct tokens in one condition. The three bars are the weighted losses (shorter is better) and the verdict compares them. Push “call, no result” up until it matches “call with result”: the difference shrinks and the call is thrown away, because the answer added nothing beyond the question. Now raise τf: stricter thresholds keep fewer, more obviously useful calls, which is the trade-off the paper tunes per tool.
Real candidates, ranked by the filter score
Ten candidate calls the paper reports (Table 10, paraphrased here), with their score L⁻ − L⁺ and the authors' own judgement of whether the call was really useful. Slide the threshold.
Reading it: rows above the line pass the filter. High scores mostly match useful calls (a question-answering lookup for the Nile's length, a calendar date that the next words repeat, a calculator ratio). Negative scores are clearly useless calls. The filter is not perfect in the middle: a search result about a band called “Fast Train” scored 0.92 despite being irrelevant, while a useful translation scored only 0.70. The paper argues a little noise is even healthy: it stops the fine-tuned model from blindly trusting every tool result.
Fine-tuning and inference · original
Everyday picture
After filtering, the kept calls are written into the original passages, and the model is fine-tuned on this annotated text with the ordinary next-token objective. Because the text is otherwise unchanged, the model learns the same language as before plus where calls help. At use time, the model writes normally until it emits the arrow “->”. Generation then pauses, the real tool runs, its result is pasted in, and generation continues.
Reading it: each step adds what the model writes, in plain text, or what the system inserts, highlighted. The key moment is the arrow: the model does not guess the result, it stops and waits for the real tool. That handshake between model and runtime is exactly how tool calling works today; see the agent loop lesson.
Two decoding tweaks mattered in practice. The model opens a call whenever “[” is among its 10 most likely next tokens (not only when it is the single most likely), which makes it far more willing to use tools. And it may make at most one call per input, so it cannot get stuck calling tools forever.
3 Tools · original
The only requirements for a tool: its input and output can be written as text, and a few examples of its use can be written down. Examples below are written for this page.
| Tool | What powers it in the paper | Example call → result |
|---|---|---|
| Question answering | Atlas, a retrieval-augmented model fine-tuned on Natural Questions | QA(Where is the Eiffel Tower?) → Paris |
| Calculator | The four basic operations, results rounded to 2 decimals | Calculator(18 * 4 + 6) → 78 |
| Wikipedia search | BM25 keyword search over a Wikipedia dump, returning short snippets | WikiSearch(Ada Lovelace) → “Augusta Ada King, Countess of Lovelace, was …” |
| Machine translation | NLLB (600M parameters), any of 200 languages into English | MT(sécurité nucléaire) → nuclear safety |
| Calendar | Returns today's date; takes no input | Calendar() → Today is Friday, March 3, 2023. |
4 Experiments · original
Setup
- Model: GPT-J, 6.7 billion parameters. Text: a subset of CCNet web text, pre-filtered per tool (for example, only passages with at least three numbers for the calculator).
- Fine-tuning: batch size 128, learning rate 1 × 10⁻⁵ with linear warmup over the first 10% of training.
- Evaluation: zero-shot. Models get an instruction but no worked examples, and must decide for themselves whether to use a tool.
- Comparisons: plain GPT-J, GPT-J fine-tuned on the same text without calls, Toolformer with calls disabled, and much larger models: OPT (66B) and GPT-3 (175B).
| Model | Facts (T-REx) | Math (ASDiv) | Questions (WebQS) | Dates (Dateset) |
|---|---|---|---|---|
| GPT-J (6.7B) | 31.9 | 7.5 | 18.5 | 3.9 |
| Toolformer, calls disabled | 34.9 | 14.8 | 18.9 | 5.9 |
| Toolformer (6.7B) | 53.5 | 40.4 | 26.3 | 27.3 |
| GPT-3 (175B) | 39.8 | 14.0 | 29.0 | 0.8 |
Reading it: each group compares the plain model, Toolformer and GPT-3 on one task (scores are the percentage answered correctly under each benchmark's lenient check). With its tools, the small model beats a model 25 times its size on facts, arithmetic and dates. On open questions, where it relies on a simple keyword search it cannot refine, it still trails GPT-3. The model also chose the right tool almost every time: the question-answering tool on 98.1% of fact questions, the calculator on 97.9% of maths problems, and search on 99.3% of open questions.
Two important non-results
- No cost to ordinary text. Perplexity on WikiText was 10.3 with or without the tool-call training (with calls disabled), so learning tools did not damage general language modelling.
- Size matters. Applying the method to smaller GPT-2 models showed that useful tool use only emerges at around 775 million parameters; smaller models score the same with and without tools.
Why it matters today
The pattern (a small model plus the right tool beating a large model alone) is the economic argument behind giving agents calculators, search and code execution, and behind routing easy work to small models; see the cost lesson.
5 Analysis · original
How eager should the model be to call?
With ordinary greedy decoding (open a call only if “[” is the single most likely token), the model called a tool on 40.3% of fact questions and only 8.5% of open questions. Allowing a call whenever “[” is in the top 10 raised that to 98.1% and 100%. Interestingly, at the strict setting the model was somewhat calibrated: it tended to call tools on exactly the questions it would otherwise get wrong.
Why it matters today
“When should the agent use a tool?” is still a tuning problem: too eager wastes time and money, too reluctant leaves errors unfixed. Clear tool descriptions and instructions play the role that the top-k trick played here; see the tools lesson.
7–8 Limitations and conclusion · original
The authors list the gaps plainly:
- No chaining. Calls were sampled independently, so the training data never shows using one tool's output as another's input (for example, get today's date, then ask a question about it).
- No interaction. The model cannot browse through search results or refine a query.
- Sensitivity to wording. Whether it calls a tool can depend on small changes in phrasing.
- Sample inefficiency. Over a million documents yielded only a few thousand useful calculator calls.
The first two are exactly what agent loops such as ReAct provide: multiple steps, each able to use earlier results.
From Toolformer to today
| In the paper | Today | Learn it |
|---|---|---|
| Calls written as text with bracket markers | Structured tool calls checked against a JSON Schema | tools |
| One call per input | Loops with many calls, often several in parallel | agent loop |
| Five fixed tools | Large, changing tool sets, loaded dynamically or over MCP | MCP |
| Self-supervised filter chooses training data | Curated tool-use training data; evaluation of tool selection and argument accuracy | evals |
Glossary
Every term with hover guidance on this page, in one place.