Chain-of-Thought Prompting, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
- The diagrams are live: hover or tap a box to read what it holds. The charts answer the arrow keys once you tab to them.
Each idea climbs the same ladder: an everyday picture, a tiny example you could check by hand, a diagram, the math, and why it still matters. The reasoning models lesson builds the mechanism from scratch in code: a toy model that can do one addition per token, and exactly how writing steps down rescues it.
Abstract
“We explore how generating a chain of thought, a series of intermediate reasoning steps, significantly improves the ability of large language models to perform complex reasoning.”Wei et al. (2022), Abstract (dashes in the original replaced with commas). Read the original
Everyday picture
A maths teacher who only ever shows the class the final answers (“question 4: 27”) gets pupils who guess. A teacher who works a few problems on the board, line by line, gets pupils who write their own working, and they get the answers right far more often. This paper does the second thing to a language model: it puts a handful of fully worked examples in the prompt, and the model imitates the working.
What the paper claims
- Showing eight worked examples, each with its intermediate steps written in plain language, makes large models dramatically better at arithmetic word problems, commonsense questions and simple symbolic puzzles. No training, no fine-tuning: only a different prompt.
- On GSM8K, a set of grade-school maths word problems, PaLM 540B went from 17.9% solved to 56.9%, beating the previous best of 55% (a fine-tuned GPT-3 with a separately trained verifier).
- The trick only works in very large models. Below roughly 100 billion parameters it does nothing or makes things worse: an emergent ability of scale.
Why it matters today
Every modern reasoning model writes a long chain of thought before it answers. This paper is where that habit was first shown to pay, and named.
1 Introduction · original
“In this paper, we combine the strengths of these two ideas in a way that avoids their limitations.”Wei et al. (2022), §1
Everyday picture
Two good ideas were lying around, each with a catch. The first: models that write out a rationale before answering do better at maths, but teaching them required thousands of hand-written worked solutions. The second: large models can pick up a task from a few examples in the prompt (in-context learning), with no training at all, but that barely helped on problems that need reasoning. The paper's move is to combine them: put a few worked solutions, rationale included, in the prompt.
Tiny example: what changes in the prompt
| Rationale training | Standard few-shot prompting | Chain-of-thought prompting | |
|---|---|---|---|
| Examples needed | thousands of worked solutions | about eight question and answer pairs | about eight question, working, answer triples |
| Model weights change? | yes: training or fine-tuning | no | no |
| Good at multi-step problems? | yes | poorly, and it barely improves with scale | yes, in large enough models |
The paper writes each example as a triple ⟨input, chain of thought, output⟩. A standard prompt uses pairs ⟨input, output⟩. That one extra element is the whole method.
Why it matters
Because nothing is trained, one model checkpoint can do many reasoning tasks, and anyone who can write eight worked examples can use it. The idea of steering a model by showing it examples comes from GPT-3's few-shot prompting; this paper changes what the examples contain.
2 Chain-of-thought prompting · original
“After Jane gives 2 flowers to her mom she has 10 … then after she gives 3 to her dad she will have 7 … so the answer is 7.”Wei et al. (2022), §2, the kind of inner working the method asks a model to write
Everyday picture
Doing a sum in your head, you keep the running total somewhere. A language model has nowhere to keep it except the text it has already written: each new token is computed by one pass through the network, and the only thing that pass can read is the text so far. Writing “23 − 20 = 3” puts the 3 on the page, where the next pass can pick it up.
Tiny example: the paper's Figure 1
Both prompts below show the model the same worked example about tennis balls, then ask a new question about apples. The only difference is whether the worked example includes its working. Hover or tap each box.
Hover or tap a box. Start with the worked example at the top of either column, then compare the two outputs at the bottom.
Reading it: the dashed frames are everything the model is given; the boxes below the arrows are what it writes. The two frames differ in exactly one box: the worked example's answer (orange on the right) spells out “2 cans of 3 is 6, 5 + 6 = 11” before stating 11 (shortened here; hover it for the full text). Everything else, including the model, is identical. Faced with the new apples question, the left model imitates its example and blurts a number, which is wrong. The right model imitates its example too, so it writes one subtraction and one addition first, then an answer that follows from them. The model was not taught to reason; it was shown what an answer looks like, and a good answer includes its working.
The math: why written steps buy computation
The paper's first reason the method should help is that a chain of thought lets a model “allocate additional computation” to problems that need more steps. Here is one way to see why. This is not an equation from the paper; it is the rule the reasoning lesson builds and tests.
In words: “the number of steps a model can do one after another equals the steps one pass can do, times the number of tokens it writes.”
With the numbers: the apples question needs two operations in a row (23 − 20, then + 6). Take an illustrative model that can do L = 1 operation per token. Blurting the number means T = 1 token of work, so S = 1: one operation short, and it must guess. Writing “23 − 20 = 3” and “3 + 6 = 9” before the answer gives T = 3, so S = 3, which is enough.
In Python:
# illustrative: one operation per forward pass
L = 1
# answer at once: one token of work
L * 1 # → 1
# two written steps, then the answer
L * 3 # → 3
# the two steps the chain writes down
step1 = 23 - 20 # → 3
step1 + 6 # → 9
The four properties the paper lists
- More computation where it is needed. A problem with more steps gets a longer chain, so more tokens, so more passes through the network.
- A window into the model. The chain shows how the model reached its answer, so you can see where it went wrong (the paper adds that this is not a full account of the model's computation).
- General. Anything a person can solve by writing steps in language is, in principle, a candidate: maths, commonsense, symbol puzzles.
- Cheap to use. A large, off-the-shelf model does it when the prompt's examples include chains.
Why it matters today
Property 1 is the seed of what is now called test-time compute: spending more tokens while answering to get a better answer. Property 2 comes with a warning the field later confirmed: a written chain is evidence about the model's reasoning, not a guaranteed record of it. In code, the lesson's solve is exactly this toy model, returning an Attempt with its steps, its answer and the tokens it spent.
3 Arithmetic reasoning · original
“Though simple for humans, arithmetic reasoning is a task where language models often struggle.”Wei et al. (2022), §3
Everyday picture
Maths word problems are a good test bench because the answer is a number that is either right or wrong, and because reaching it takes several dependent steps: read the story, pick out the quantities, decide the operations, do them in order.
3.1 Experimental setup · original
Everyday picture
Same exam, two ways of briefing the candidate. Every model gets the same eight example problems; in one briefing the examples show only answers, in the other they show the working too. Then the model sits five sets of word problems.
Tiny example: one of the eight worked examples
| Q | There were nine computers in the server room. Five more computers were installed each day, from monday to thursday. How many computers are now in the server room? |
| A | There were originally 9 computers. For each of 4 days, 5 more computers were added. So 5 * 4 = 20 computers were added. 9 + 20 is 29. The answer is 29. |
What was tested
- Benchmarks: GSM8K, SVAMP, ASDiv, AQuA (multiple choice) and MAWPS, five collections of maths word problems.
- Prompts: one hand-written set of eight exemplars for every benchmark except AQuA, which used four from its own training set. The authors note the exemplars were not tuned.
- Models: GPT-3 (four sizes, 350M to 175B), LaMDA (422M to 137B), PaLM (8B, 62B, 540B), UL2 20B and Codex.
- Decoding: greedy: always the single most likely next token, so one chain per question. (The paper points to follow-up work that samples many chains and votes; that is the self-consistency companion.)
Why it matters
The comparison is clean: same model, same questions, same number of examples. The only variable is whether the examples show their working, so any difference in accuracy is caused by the chains.
3.2 Results · original
“Chain-of-thought prompting does not positively impact performance for small models, and only yields performance gains when used with models of ∼100B parameters.”Wei et al. (2022), §3.2
Everyday picture
Show a toddler a worked long-division example and they will copy its shape: numbers, lines, a confident final figure, all nonsense. Show it to a ten-year-old and they will actually follow it. Small models are the toddler: the paper found they wrote “fluent but illogical” chains, and did worse than if they had just answered.
Tiny example
On GSM8K, the smallest LaMDA (422M) solved 2.6% with standard prompting and 0.4% with chains: worse. LaMDA 137B went from 6.5% to 14.3%, and PaLM 540B from 17.9% to 56.9%, more than three times as many problems solved.
The paper's Figure 4, redrawn
Hover the chart, or tab to it and use the arrow keys, to read each model size.
Reading it: the x-axis is model size in billions of parameters, on a log scale (each labelled tick is ten times the one before); the y-axis is the share of problems solved, from 0 to 100%. The solid line is standard prompting, the dashed line chain-of-thought prompting, and the dotted flat line the best result any fine-tuned system had reached before. Start with PaLM on GSM8K: at 8B the two lines sit together near 5%, then the dashed line takes off, passing the dotted line at 540B. Switch to LaMDA or GPT and the small models' dashed line sits below the solid one: chains hurt until the model is big enough. That sudden bend, flat and then steep, is what the paper means by emergent. Now try GPT on SVAMP: standard prompting already reaches 65.7% at 175B and chains add only 3.2 points, while on GSM8K, the hardest set, the same model gains 31.3. Numbers from Table 2 of Wei et al. (2022), reproduced under CC BY 4.0.
The math: the gain, and when it turns positive
In words: “the gain from chain-of-thought prompting, for a model of a given size, is its accuracy with chains minus its accuracy without.”
With the numbers: on GSM8K, LaMDA 422M: Δ = 0.4 − 2.6 = −2.2 points (chains hurt). LaMDA 137B: 14.3 − 6.5 = +7.8. PaLM 540B: 56.9 − 17.9 = +39.0, and 56.9 / 17.9 = 3.18 times as many problems solved.
In Python:
# GSM8K accuracy (%), standard then chain of thought, from Table 2
acc = {"LaMDA 422M": (2.6, 0.4), "LaMDA 137B": (6.5, 14.3), "PaLM 540B": (17.9, 56.9)}
# Δ(N) = acc_CoT(N) − acc_std(N)
{N: round(cot - std, 1) for N, (std, cot) in acc.items()} # → {'LaMDA 422M': -2.2, 'LaMDA 137B': 7.8, 'PaLM 540B': 39.0}
# how many times as many problems PaLM 540B solves with chains
round(56.9 / 17.9, 2) # → 3.18
Three takeaways the paper draws
- It is emergent. Gains appear only around 100B parameters; smaller models write chains that read well and reason badly.
- Harder problems gain more. GSM8K, the hardest set, more than doubled for the largest GPT and PaLM models. On SingleOp, the one-step subset of MAWPS, the gain was tiny or negative (PaLM 540B: 94.1% either way).
- It rivals fine-tuning. PaLM 540B with eight exemplars set new records on GSM8K, SVAMP and MAWPS, against systems trained on each dataset's training set.
Why it matters today
The shape of this chart, a method useless for small models and powerful for large ones, is why reasoning tricks are usually tested on the largest models first, and why training small models to reason later relied on distillation and reinforcement learning rather than prompting alone.
3.3 Ablation study · original
Everyday picture
When a new medicine works, you want to know which ingredient did it. An ablation removes one ingredient at a time. The paper tests three rival explanations for why chains help, each by building a prompt that keeps that one ingredient and drops the rest.
Tiny example: the same exemplar, four ways
| Variant | The exemplar's answer | Tests the idea that… |
|---|---|---|
| Equation only | 5 * 4 + 9 = 29. The answer is 29. | the help comes from writing the maths |
| Variable compute only | a row of dots, one per character of that equation, then the answer | the help comes from extra tokens, whatever they say |
| Reasoning after answer | The answer is 29. There were originally 9 computers… | chains merely remind the model of what it knows |
| Chain of thought | There were originally 9 computers… The answer is 29. | (the full method) |
Reading it: each bar is the share of GSM8K problems LaMDA 137B solved, on an axis from 0 to 20%. The three ablations (grey) all land at or below standard prompting's 6.5%; only the full chain of thought (blue) jumps, to 14.3%. So it is not the equation alone (GSM8K's stories are too tangled to turn straight into one equation), not the extra tokens alone (dots add passes but carry no intermediate results), and not a reminder effect (a chain written after the answer cannot change the answer). Numbers from Table 6 of Wei et al. (2022), reproduced under CC BY 4.0.
The math: why the order matters
A language model writes left to right, and each piece is predicted from everything before it. So the chance of writing a chain and then an answer splits into two factors, and so does the reverse order.
And with the answer written first:
In words: “the chance of writing this chain and then this answer is the chance of the chain, times the chance of the answer once the chain is on the page. Put the answer first and its chance is fixed before any working exists; the chain written afterwards cannot change it.”
With the numbers (illustrative, not from the paper): for the apples question, say the model writes the chain “23 − 20 = 3. 3 + 6 = 9.” with probability 0.40. With 3 and 6 on the page, writing 9 next is nearly a copy: probability 0.98. The pair together: 0.40 × 0.98 = 0.392. Answering first, the model must produce 9 with no working, say with probability 0.10, and that 0.10 is the accuracy no matter how good the explanation that follows.
In Python:
# illustrative probabilities for the apples question
# P(c | E, x): the model writes this chain
p_c = 0.40
# P(y | E, x, c): with 3 and 6 written, 9 is nearly a copy
p_y_given_c = 0.98
# P(c, y | E, x)
round(p_c * p_y_given_c, 3) # → 0.392
# answer first: P(y | E, x) is settled before any working is written
p_y = 0.10
p_y # → 0.1
Adding up the first product over every chain that ends in 9 gives the model's full chance of answering 9. Greedy decoding follows one chain only; the self-consistency paper samples many and lets their answers vote, which estimates that sum.
Why it matters today
The dots result is subtle and still debated. Extra tokens are extra computation, as §2 argued, but the dots carry nothing from one pass to the next: the benefit needs the intermediate results to be written down where the next pass can read them. The lesson's toy model makes the same point: its running total survives only because it is written into the text.
3.4 Robustness of chain of thought · original
Everyday picture
If a recipe only works when one particular chef writes it, it is a trick, not a method. The paper had two more authors write their own chains for the same eight questions, wrote a deliberately terse version, and also borrowed three random sets of eight worked solutions written by crowd workers for GSM8K's training set.
Tiny example: the concise style
The original chain for the computers question uses several short sentences (“For each of 4 days, 5 more computers were added. So 5 * 4 = 20 computers were added. 9 + 20 is 29.”). The concise one reads “5 * 4 = 20 new computers were added. So there are 9 + 20 = 29 new computers in the server room now”.
Reading it: each bar is LaMDA 137B's GSM8K solve rate, axis 0 to 20%. The top grey bar is standard prompting (6.5%). Every chain-of-thought variant beneath it, whoever wrote it and in whatever style, lands between 11.1% and 17.6%: 1.7 to 2.7 times standard prompting. The spread between them is real (prompts matter), but the gap to the grey bar is bigger than the spread. Numbers from Table 6 of Wei et al. (2022), reproduced under CC BY 4.0.
The appendix adds that results held across different exemplar orders and different numbers of exemplars, with one warning: on the coin-flip task, annotator A's chains scored 99.6% and annotator C's 71.4% (standard prompting: 50.0%). Prompt writing still matters.
Why it matters today
This is why “show worked examples with steps” became general advice rather than one lab's trick, and why later work could generate the chains automatically instead of writing them by hand.
4 Commonsense reasoning · original
Everyday picture
Not every hard question is maths. “Would a pear sink in water?” needs a fact (a pear's density) and one inference (less dense than water floats). Chains of thought are written in ordinary language, so they apply to reasoning about the everyday world as naturally as to sums.
Tiny example: the five tasks
- CSQA: multiple-choice commonsense questions.
- StrategyQA: yes/no questions that need an unstated multi-step strategy.
- Date understanding: work out a date from a short story.
- Sports understanding: is a sentence about sport plausible?
- SayCan: turn an instruction into a sequence of robot actions.
Reading it: each pair of bars is one task for PaLM 540B, axis 0 to 100%: plain grey for standard prompting, striped blue for chain of thought. Every task improves, but by very different amounts: Date (+16.3 points) and Sports (+14.9) gain a lot, CSQA barely moves (+1.8). The paper highlights two results: 75.6% on StrategyQA beat the prior best of 69.4%, and 95.4% on Sports beat an unaided sports enthusiast (84%). Numbers from Table 4 of Wei et al. (2022), reproduced under CC BY 4.0. (The main text reports StrategyQA as 75.6%; Table 4 lists 77.8%.)
Why it matters
Where the question needs one fact rather than several steps (much of CSQA), a chain has little to do. Where the answer comes from combining several facts or small inferences, it helps. The same rule, “chains help where there are steps”, reappears in the appendix's advice below.
5 Symbolic reasoning · original
Everyday picture
Two puzzles a child can do with a pencil: take the last letter of each word in a name and join them (“Amy Brown” becomes “yn”); and track a coin as people flip it or leave it alone. They are easy to generate in any length, which makes them a clean test of a harder question: shown examples with two steps, can the model handle three or four?
Tiny example: track the coin yourself
Reading it: press a person's button to switch between “flips” and “does not flip”. The question is written the way the paper's puzzles are; the orange chain below it follows the pattern of the paper's eight coin-flip exemplars (Table 23), which all have exactly two people. The chain reduces the story to one count and one rule: an even number of flips leaves the coin heads up. Add a third or fourth person and you leave the exemplars' length: the readout shows how PaLM 540B did at that length with and without chains.
The math: the rule the chain uses
In words: “count the people who flip; the coin is still heads up exactly when that count is even.”
With the numbers: the paper's example: “Phoebe flips the coin. Osvaldo does not flip the coin.” So f = (1, 0), the sum is 1, and 1 mod 2 = 1, which is not 0: the coin is tails up, and the answer is “no”.
In Python:
# Phoebe flips, Osvaldo does not
f = [1, 0]
# Σ f_i: how many flips
sum(f) # → 1
# heads up if the count is even
sum(f) % 2 == 0 # → False
# the other puzzle: last letters of "Amy Brown"
"".join(word[-1] for word in "Amy Brown".split()) # → 'yn'
Reading it: PaLM 540B's solve rate, axis 0 to 100%, plain grey for standard prompting and striped blue for chain of thought. “2” is the exemplars' own length (in domain); “3” and “4” are longer than anything in the prompt (out of distribution). Standard prompting never manages last-letter concatenation at any length (7.6% at best) and falls to a coin toss (about 50%) on longer coin puzzles. With chains, last letters go from 99.4% at two words to 63.0% at four, and coins stay above 90%. The chains taught a procedure that stretches beyond the examples' length, imperfectly. Numbers from Table 5 of Wei et al. (2022), reproduced under CC BY 4.0.
Why it matters
The paper stresses that in the in-domain case the exemplars hand the model the exact procedure; it only has to repeat it with new names. Small models still fail at that: in this paper, the ability to carry a procedure over to unseen symbols appeared only around 100B parameters.
6 Discussion and limitations · original
“Standard prompting only provides a lower bound on the capabilities of large language models.”Wei et al. (2022), §6
Everyday picture
Judging a person's maths by making them answer instantly, with no paper, measures their mental arithmetic, not their maths. The paper's point is the same about models: how you ask changes what you measure, so a flat curve under one prompting style does not mean the model cannot do the task.
The four limitations the authors list
- Imitating the look of human reasoning does not settle whether the network is “reasoning”; they leave that open.
- Writing eight chains is cheap, but writing thousands for fine-tuning would not be (they suggest synthetic data).
- Nothing guarantees a correct chain: wrong chains can reach right answers and vice versa.
- It only works in very large models, which are expensive to serve.
Why it matters today
All four turned into research programmes. Limitation 2 was answered by generating chains with models and training on the ones that reach correct answers; limitation 3 by verifiers that check chains step by step; limitation 4 by distilling reasoning into smaller models. The reasoning lesson covers verifiers and the reinforcement learning that now trains the habit in directly.
Appendices: why scale, and where chains go wrong · original
Everyday picture
Marking a pile of wrong exam scripts tells you more than the score does: was it a slip on the calculator, a misread question, or a step left out? The authors hand-graded model chains the same way.
Tiny example: 50 wrong chains, graded
| What was wrong | Example from the paper | Share |
|---|---|---|
| Calculator error only | “3 x 25 x 8 = 300” (it is 600) | 8% |
| Symbol mapping error | multiplied a coach's 15 weekly hours by 30 (the pay rate) instead of 50 (the weeks) | 16% |
| One step missing | found the second recipe's 40 instructions, forgot to add the first recipe's 20 | 22% |
| Major errors: misunderstood the question, or incoherent | subtracted 30 from the total of 110 coins and called the result the silver coins | 54% |
So nearly half the wrong chains (8 + 16 + 22 = 46%) were almost right. Among 50 chains that reached a correct answer, the logic was sound in all but one or two (§3.2 counts two coincidences; Appendix D.1 lists one). The paper cautions that this is easier to trust for numeric answers than for yes/no or multiple choice, where a wrong chain can land on the right option by luck.
Two more findings. Passing every equation in a chain through a real calculator (Python's eval) lifted LaMDA 137B's GSM8K score from 14.3% to 17.8% and PaLM 540B's from 56.9% to 58.6% (Table 1). And grading 45 errors from PaLM 62B (20 semantic misunderstandings, 18 missing steps, 7 other) showed that scaling to PaLM 540B fixed a substantial share in every category (Appendix A.1).
When to expect a gain (Appendix A.3)
The authors' rule of thumb: chains help most when (1) the task needs several steps, (2) the model is large, and (3) standard prompting's accuracy grows little with scale. Remove any one and the gain shrinks.
Why it matters today
The calculator result foreshadowed tool use: let the model plan in words and hand the arithmetic to a program. The error categories foreshadowed process reward models, which score each step of a chain rather than only its answer.
7–8 Related work and conclusion · original
The paper places itself between two lines of work. One trained or fine-tuned models to write intermediate steps: natural-language rationales for maths (Ling et al., 2017; Cobbe et al., 2021) and line-by-line “scratchpads” for program execution (Nye et al., 2021). The other improved prompting, mostly by changing the input side (instructions, learned prompts). Chain-of-thought prompting instead changes the output the examples demonstrate, and needs no training. The conclusion restates the main finding: for tasks with flat scaling curves under standard prompting, chains of thought produce steeply rising ones, in large enough models.
What changed since 2022
| In the paper | Today | Learn it |
|---|---|---|
| Eight hand-written worked examples in the prompt | Models trained to reason; often a plain instruction, or nothing, is enough to start a chain | reasoning |
| One greedy chain per question | Many sampled chains and a vote, or a verifier picking the best | self-consistency companion |
| Chains a few sentences long | Thousands of thinking tokens, with a budget the caller sets | reasoning |
| Hand-graded error analysis | Automatic checks: unit tests, exact-match answers, step-level verifiers | reasoning |
| An external calculator bolted on afterwards | Tool calls in the middle of reasoning | ReAct companion |
Glossary
Every term with hover guidance on this page, in one place.