An annotated companion · AI Primer

ReAct: reasoning and acting, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is released under a Creative Commons Attribution 4.0 licence, so selected tables and one of its prompt examples are reproduced here with attribution to Yao et al. (2022). Figures are redrawn from scratch. Read the original alongside: every section links to it.

How to read this page

  • Any dotted word explains itself when you hover, tab to, or tap it, and so does every symbol in every equation.
  • The step-through replays one question four ways: answering directly, reasoning only, acting only, and ReAct.

Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. The production version of this loop, with budgets and loop detection, is in the agent loop lesson.

Abstract

“In this paper, we explore the use of LLMs to generate both reasoning traces and task-specific actions in an interleaved manner, allowing for greater synergy between the two.”Yao et al. (2022), Abstract. Read the original

Everyday picture

A detective who only thinks, never checking facts, builds elegant theories on false premises. A detective who only runs around gathering evidence, never stopping to think, collects clues without knowing what they mean. A good detective alternates: think about what to check, check it, think about what was found, decide the next step. ReAct asks a language model to work like the good detective, writing its thinking and its actions in one interleaved transcript.

What the paper claims

  • Interleaving reasoning traces with actions beats doing either alone on interactive tasks, and on fact verification.
  • Reasoning grounded in fetched facts suffers far less from hallucination than reasoning from memory alone.
  • With one or two examples in the prompt, it beat imitation and reinforcement learning systems trained on thousands of examples: +34 points of success rate on ALFWorld and +10 on WebShop.
  • The transcripts are readable, so people can inspect and even correct the agent mid-task.

Why it matters today

The think, act, observe loop in this paper is the skeleton of nearly every agent built since. Modern APIs formalise the “Action” line as structured tool calls, but the shape of the loop is ReAct's.

1 Introduction · original

“I don’t have salt, so let me use soy sauce and pepper instead.”Yao et al. (2022), §1, an example of the inner speech that guides human action

Everyday picture

The paper's own example is cooking. Between actions you talk to yourself: to track progress (“everything is cut, now heat the water”), to handle surprises (the salt example above), and to notice when you need information (“how do I make dough? let me look it up”). And you act to support your thinking: open the cookbook, check the fridge.

The two traditions it joins

Reason only (chain of thought)Act only
What it doesWrites step-by-step reasoning, then answersIssues actions (search, click, move) and reads the results
StrengthPlans, decomposes, does arithmeticGets real, current information
WeaknessA “static black box”: facts come from memory and can be invented; errors propagateNo working memory or plan: loses track, repeats itself, fails to combine what it found

ReAct combines them: reason to act (plans guide which action to take) and act to reason (observations feed the next thought).

2 The idea: thoughts as actions · original

Everyday picture

Picture an agent's options as a menu of buttons: “search”, “look up”, “finish”. ReAct adds one more button labelled “think”. Pressing it changes nothing in the outside world and returns no new information. It only writes a note on the agent's own notepad, and that note is visible when choosing the next button.

Context c_tquestion + history Thoughtin language Actionsearch / lookup / finish EnvironmentWikipedia API Observationtool result

Hover or tap a block, starting with Context at the top.

The ReAct loop, drawn for this page from the method in §2 of Yao et al. (2022). The dashed arrow is a thought going straight back into the context.

Reading it: start at the context: the question plus everything written so far. The model writes a thought. The dashed arrow shows the thought going straight back into the context, because it touches nothing outside. Then it writes an action, which does go outside: the environment runs it and returns an observation, which is appended to the context too. Round and round, until the action is “finish”. The only difference from an act-only agent is the thought box; the only difference from chain-of-thought is the environment box.

The math

In words: “the agent chooses each action from everything it has seen and done so far. ReAct enlarges the menu of actions with every possible sentence; choosing a sentence (a thought) gets no reply from the world and simply adds the sentence to the context.”

With the numbers: in the example below, after one round the context is c = (question, Thought 1, Search[Colorado orogeny], its observation). Thought 2, “It does not mention the eastern sector”, is an element of 𝓛; it produces no observation, only c = (…, Thought 2), and that is what makes Action 2, Lookup[eastern sector], an obvious next move.

In Python:

# c_t
c = ["question", "Thought 1", "Search[Colorado orogeny]", "its observation"]
# actions from 𝒜
A = ["Search[Colorado orogeny]", "Lookup[eastern sector]"]
# â_t, a sentence from 𝓛
a_hat = "Thought 2: It does not mention the eastern sector"
# not an action the world answers ...
a_hat in A  # → False
# ... so c_(t+1) = (c_t, â_t), with no observation
c = c + [a_hat]
len(c), c[-1]  # → (5, 'Thought 2: It does not mention the eastern sector')

Why it matters today

This framing, a policy choosing from actions given a growing history, is how agent frameworks are built. The history is the message list; the actions are tools; the “thoughts” today are often the model's visible reasoning between tool calls. See talking to a model for the message format.

Try it: four ways to answer

The question: What is the elevation range for the area that the eastern sector of the Colorado orogeny extends into? The ReAct, Act and Reason-only traces below are the paper's own few-shot examples (Appendix C, reproduced under CC BY 4.0), the demonstrations the model imitates. The two failure traces are illustrative, written for this page to show the failure modes the paper reports in §3.3.

Reading it: step through the ReAct trace first. Thoughts (blue) plan and interpret; actions (orange) touch Wikipedia; observations (green) are what came back. Notice Thought 4: the search for “High Plains” returned a disambiguation page, and the thought is what decides to retry with “High Plains (United States)”. The Act-only example makes the same moves only because a human wrote it; in the illustrative failure, with no thought to interpret the ambiguous result, the agent repeats itself. The Reason-only example is correct because the fact is well known; in its illustrative failure, the model confidently recalls a wrong range and nothing checks it.

3 Knowledge-intensive reasoning · original

3.1–3.2 Setup and methods · original

Everyday picture

The agent gets a deliberately clumsy browser: it can open a Wikipedia page by exact title (seeing only its first five sentences), search within the open page like Ctrl+F, or submit an answer. Clumsy on purpose, so success depends on reasoning about what to look up next, as a person would.

ActionWhat it returns
search[entity]The first 5 sentences of that entity's Wikipedia page, or the top-5 similar titles if there is no exact match
lookup[string]The next sentence on the current page containing the string
finish[answer]Ends the task with that answer
  • Tasks: HotpotQA (questions needing two or more Wikipedia passages) and FEVER (claims to label as supported, refuted, or not enough information). The model sees only the question or claim.
  • Model: PaLM-540B, frozen, prompted with 6 (HotpotQA) or 3 (FEVER) hand-written example trajectories. Adding more examples did not help.
  • Baselines: the same examples with parts removed: answers only (Standard), thoughts only (CoT), actions only (Act), plus self-consistency (CoT-SC: sample 21 reasoning chains at temperature 0.7 and take the majority answer).

Combining internal and external knowledge

ReAct was more factual; chain of thought was better at structuring reasoning. So the paper lets each fall back to the other:

  • ReAct → CoT-SC: if ReAct has not answered within the step limit (7 steps for HotpotQA, 5 for FEVER), switch to self-consistent chain of thought.
  • CoT-SC → ReAct: if the reasoning samples do not agree strongly enough, switch to ReAct and go look things up:

In words: “if fewer than half of the sampled reasoning chains agree on the most popular answer, the model's own knowledge is shaky, so go and look it up.”

With the numbers: with n = 21 samples, n/2 = 10.5. If the most common answer appears 9 times, 9 < 10.5, so switch to ReAct. If it appears 14 times, trust it.

In Python:

# sampled reasoning chains
n = 21
n / 2  # → 10.5
# votes for the most common answer
c_maj = 9
# back off to ReAct?
c_maj < n / 2  # → True
c_maj = 14
c_maj < n / 2  # → False

Fine-tuning

The paper also generated 3,000 successful ReAct trajectories and fine-tuned smaller PaLM models (8B and 62B) on them.

3.3 Results · original

Table 1 of Yao et al. (2022), reproduced under CC BY 4.0: PaLM-540B prompting results. HotpotQA exact match; FEVER accuracy.
Prompt methodHotpotQA (EM)FEVER (Acc)
Standard28.757.1
CoT29.456.3
CoT-SC33.460.4
Act25.758.9
ReAct27.460.9
CoT-SC → ReAct34.264.6
ReAct → CoT-SC35.162.0
Supervised state of the art67.589.5

Reading it: compare rows in pairs. ReAct beats Act on both tasks, so reasoning helps acting. ReAct beats CoT on FEVER (where tiny factual differences decide the label) but trails it slightly on HotpotQA. The two combined methods at the bottom are best on each task, which is the paper's central message: internal reasoning and external lookup cover each other's weaknesses. All prompting methods remain far below the supervised systems trained specifically for these benchmarks.

How each method fails

Table 2 of Yao et al. (2022), reproduced under CC BY 4.0: success and failure modes in 200 human-labelled HotpotQA trajectories (50 correct and 50 incorrect from each method).
TypeDefinitionReActCoT
Success: true positiveCorrect reasoning trace and facts94%86%
Success: false positiveHallucinated reasoning trace or facts6%14%
Failure: reasoning errorWrong reasoning trace, including failing to recover from repetitive steps47%16%
Failure: search result errorSearch returned nothing useful23%–
Failure: hallucinationHallucinated reasoning trace or facts0%56%
Failure: label ambiguityRight prediction that did not match the label exactly29%28%

Reading it: the highlighted row is the headline. More than half of chain of thought's failures were invented facts; ReAct had none, because its facts came from observations. ReAct's own weak spots are different: reasoning errors (47%), including a characteristic loop where it repeats earlier thoughts and actions, and bad search results derailing it (23%). Those two are exactly what modern agent loops defend against with loop detection and better retrieval.

Fine-tuning flips the ranking

With small models and prompting only, ReAct was the worst method: learning to reason and act from a few examples is hard. After fine-tuning on 3,000 trajectories it became the best: fine-tuned PaLM-8B ReAct beat every PaLM-62B prompting method, and fine-tuned 62B beat every 540B prompting method. Fine-tuning on answers teaches a model to memorise facts; fine-tuning on ReAct trajectories teaches it how to go and find them.

Why it matters today

The failure table is an evaluation template. For any agent, label a sample of runs by failure mode (wrong tool, bad retrieval, loop, invented fact) before trying to fix anything; see the evaluation lesson and the failure catalogue.

4 Decision-making tasks · original

Everyday picture

Two text games with long horizons. In ALFWorld, a simulated household, the agent receives goals like “examine the paper under the desk lamp” and must move between up to 50 locations, open things and use objects, sometimes over 50 steps. In WebShop, a shopping site with 1.18 million real products, it must buy an item matching instructions such as “a nightstand with drawers, nickel finish, under $140”. Here thoughts are sparse: the model decides for itself when a thought is worth writing.

From Tables 3 and 4 of Yao et al. (2022), reproduced under CC BY 4.0: ALFWorld overall success rate (%) and WebShop score and success rate.
ALFWorld methodSuccess (%)WebShop methodScoreSuccess (%)
BUTLER (imitation learning, best of 8)37Imitation learning59.929.1
Act (best of 6 prompts)45Imitation + reinforcement learning62.428.7
ReAct-IM (dense feedback only, best of 6)53Act62.330.1
ReAct (average of 6)57ReAct66.640.0
ReAct (best of 6)71Human expert82.159.6

Reading it: each bar is overall ALFWorld success on the 134 unseen games. ReAct's best prompt reaches 71%, against 45% for the best act-only prompt and 37% for BUTLER, which was trained on example trajectories. Even ReAct's worst prompt (48%) beat the best of both. The ReAct-IM bar is an ablation whose thoughts only restate what the environment said (“inner monologue”); it lands in between, which shows the value is in genuine reasoning (decomposing goals, using common sense about where desk lamps usually are), not just in narrating feedback.

Why it matters today

Without thoughts, the act-only agent loses track of what it has already done and what subgoal comes next. Production agents solve the same problem more explicitly: a written plan, task state kept outside the model, and progress checks. See the planning lesson.

5–6 Related work and conclusion · original

The paper places ReAct next to chain-of-thought reasoning (and its variants such as self-consistency), and next to language models used as policies for web browsing (WebGPT), robots (SayCan) and closed-loop agents (Inner Monologue). Its distinguishing claim is that the reasoning is free-form and sparse and is used to plan, track and adjust, rather than only to restate observations. The conclusion notes the main limitation of prompting: complex tasks need more demonstrations than fit in a prompt, which motivates fine-tuning and, eventually, reinforcement learning on agent trajectories.

From ReAct to today's agents

In the paperTodayLearn it
“Action: search[…]” parsed from free textStructured tool calls validated against a JSON Schematools
Thoughts written as “Thought:” linesModel reasoning between tool calls, sometimes hidden or summarised by the APItalking to a model
Fixed step limit (7 or 5)Step, token and cost budgets, plus loop detectionagent loop
Back off between methodsRouting, escalation and hand-off to a humanorchestration
One Wikipedia APIMany tools, often discovered at runtime (for example over MCP)MCP

Related companions: Toolformer teaches a model when to call tools by itself, and Reflexion adds a memory of lessons from failed attempts on top of a ReAct agent.

Glossary

Every term with hover guidance on this page, in one place.