Self-Consistency, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
- The vote explorer lets you set how often one sampled answer is right and how scattered the wrong ones are, and shows what voting over 1 to 40 samples does.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. This paper builds directly on chain-of-thought prompting; read that companion first if “chain of thought” is new. The reasoning models lesson implements the vote and measures when it fails.
Abstract
“Self-consistency leverages the intuition that a complex reasoning problem typically admits multiple different ways of thinking leading to its unique correct answer.”Wang et al. (2022), Abstract. Read the original
Everyday picture
Unsure of a sum, you work it out three different ways. If two routes land on 18 and one on 26, you trust 18. Different correct routes meet at the same place; mistakes tend to scatter. This paper does that to a language model: ask it the same question many times, let it reason differently each time, and keep the answer it reaches most often.
What the paper claims
- Replacing greedy decoding (one chain of thought) with self-consistency (many sampled chains and a vote on their answers) raises accuracy on every model and task tested.
- Headline gains: GSM8K +17.9 points, SVAMP +11.0, AQuA +12.2, StrategyQA +6.4, ARC-challenge +3.9.
- It needs no training, no extra model and no labels: one off-the-shelf model, a few-shot prompt, and sampling.
Why it matters today
“Sample several attempts and pick the most common answer” is now a standard way to spend test-time compute, and a baseline every fancier method (verifiers, search, trained reasoning models) is measured against.
1 Introduction · original
“The more that deliberate thinking and analysis is required for a problem, the greater the diversity of reasoning paths that can recover the answer.”Wang et al. (2022), §1
Everyday picture
Chain-of-thought prompting, as first published, asked the model for its single most likely chain: at every step, take the most probable next token. That is like asking one person for one attempt and accepting it. Self-consistency asks the same person for many attempts, lets each wander a little differently, and trusts the answer the attempts agree on.
Tiny example: the paper's Figure 1
The question: Janet's ducks lay 16 eggs per day. She eats three for breakfast every morning and bakes muffins for her friends every day with four. She sells the remainder for $2 per egg. How much does she make every day? The right answer is (16 − 3 − 4) × 2 = 18 dollars. Hover or tap each box.
Hover or tap a box. Start with the greedy path, then the three sampled ones, then the vote.
Reading it: read from the top. The prompt and the model are the same for both methods. Greedy decoding follows the single most likely path, which confuses eggs used with eggs sold and answers 14 dollars (pink, ✗). Self-consistency samples three paths instead. Two take different routes (subtracting both at once, or one at a time) and both reach 18 dollars (green, ✓); one sets up the right expression but gets the arithmetic wrong and says 26. Only the answers, in the right-hand frame, go to the vote; the reasoning is thrown away. The two 18-dollar boxes light up together when you hover either: that shared answer is what the vote counts.
Why it matters
The paper calls it a “self-ensemble”: the benefit of asking many independent solvers, from one model. It needs nothing that chain-of-thought prompting did not already have, which is why it spread so quickly.
2 Self-consistency over diverse reasoning paths · original
“We hypothesize that correct reasoning processes, even if they are diverse, tend to have greater agreement in their final answer than incorrect processes.”Wang et al. (2022), §2
Everyday picture
A hiking group splits up to find the lake. Those who find it all end up at the same shore, whatever path they took; those who get lost end up in many different places. Count where people end up, and the biggest crowd is probably at the lake.
Tiny example: three steps
- Prompt the model with worked chain-of-thought examples, exactly as in the chain-of-thought paper.
- Sample m outputs instead of taking the greedy one: at each token, draw from the model's probabilities (with a temperature, and often top-k) so each run can take a different route. Each output is a reasoning path ri followed by “The answer is …”, from which a small parser pulls the answer ai.
- Vote: throw the paths away and return the answer that appears most often.
In Figure 1 the sampled answers are a1 = 18, a2 = 26, a3 = 18. 18 gets two votes and 26 gets one, so the answer is 18.
The math: the majority vote
In words: “for every candidate answer, count how many of the m sampled paths ended in it; return the answer with the highest count.”
With the numbers: m = 3, answers (18, 26, 18). For a = 18 the sum is 1 + 0 + 1 = 2; for a = 26 it is 0 + 1 + 0 = 1. The largest count is 2, at a = 18, so â = 18.
In Python:
# a_i: the answers parsed from m = 3 sampled paths
answers = [18, 26, 18]
m = len(answers)
# Σ_i 1(a_i = a) for every candidate answer a
counts = {a: sum(1 for a_i in answers if a_i == a) for a in sorted(set(answers))}
counts # → {18: 2, 26: 1}
# argmax over a
max(counts, key=counts.get) # → 18
The paper calls this marginalizing out the reasoning paths: the model's chance of giving an answer is the sum of its chances over every path that leads there, and counting sampled answers estimates that sum. The path is a means to the answer, not something to be graded.
Why it matters
The vote only works when two answers can be compared exactly: a number, a multiple-choice letter, yes or no. The paper says so directly, and suggests that open-ended text would need some measure of whether two answers agree. In code, this is majority_vote in the reasoning lesson: one line with a counter.
Other ways to aggregate: weighting by probability · original
Everyday picture
Should a confident friend's vote count more than a hesitant one's? Only if confidence tracks being right. The paper tried weighting each path by how probable the model found it, and found the model's confidence too flat to be worth much.
Tiny example: how probable is a path?
The chance of a whole path is the product of its tokens' chances, which shrinks with every token, so long careful chains would always lose to short ones. The fix is to average per token instead. Take an illustrative 4-token path whose tokens had probabilities 0.9, 0.5, 0.8 and 0.6: the product is 0.216, and the per-token (geometric) average is 0.682. Below, c is short for the prompt and the question.
In words: “a path's score is the typical probability of one of its tokens: take the log of each token's probability, average the logs over the path's length, and undo the log.”
With the numbers: for (0.9, 0.5, 0.8, 0.6), K = 4, the logs are −0.105, −0.693, −0.223, −0.511; their average is −0.383, and exp(−0.383) = 0.682. A 4-token path and an 8-token path whose tokens are all 0.9 have products 0.656 and 0.430 (the longer one looks worse), but both score exactly 0.9 once averaged per token.
In Python:
import math
# illustrative P(t_k | ...) for each token of one path
p = [0.9, 0.5, 0.8, 0.6]
K = len(p)
# the unnormalized path probability: the product
round(math.prod(p), 3) # → 0.216
# log P(t_k | ...) for each token
[round(math.log(p_k), 3) for p_k in p] # → [-0.105, -0.693, -0.223, -0.511]
# exp( (1/K) Σ_k log P(t_k | ...) )
round(math.exp(sum(math.log(p_k) for p_k in p) / K), 3) # → 0.682
# length: products punish long paths, the per-token average does not
round(0.9 ** 4, 3), round(0.9 ** 8, 3) # → (0.656, 0.43)
round(math.exp(sum(math.log(0.9) for _ in range(8)) / 8), 3) # → 0.9
Summing the weights, or averaging them
In words: “an answer's score is the sum of the probabilities of the paths that reached it. The ‘weighted average’ variant divides that by how many paths reached it.”
With the numbers (illustrative): give Figure 1's paths per-token scores 0.62 ($18), 0.66 ($26) and 0.60 ($18), close together as the paper observed. The weighted sum gives 18 a score of 1.22 and 26 a score of 0.66: 18 wins, the same as the plain vote. The weighted average gives 18 an average of 0.61 and 26 an average of 0.66: 26 wins, because averaging throws away how many paths agreed.
In Python:
# (a_i, per-token path probability), illustrative
paths = [(18, 0.62), (26, 0.66), (18, 0.60)]
# weighted sum: Σ_i 1(a_i = a) · P(r_i, a_i | ...)
wsum = {a: round(sum(w for a_i, w in paths if a_i == a), 2) for a in (18, 26)}
wsum # → {18: 1.22, 26: 0.66}
# weighted average: the sum divided by the number of paths that reached a
wavg = {a: round(wsum[a] / sum(1 for a_i, _ in paths if a_i == a), 2) for a in (18, 26)}
wavg # → {18: 0.61, 26: 0.66}
| Aggregation | GSM8K | AQuA | ARC-c |
|---|---|---|---|
| Greedy decode (one path) | 56.5 | 35.8 | 85.2 |
| Weighted average, unnormalized | 56.3 | 35.8 | 82.3 |
| Weighted average, normalized | 22.1 | 15.7 | 51.7 |
| Weighted sum, unnormalized | 59.9 | 38.2 | 83.5 |
| Weighted sum, normalized | 74.1 | 48.0 | 88.7 |
| Unweighted sum (majority vote) | 74.4 | 48.3 | 88.7 |
Reading it: compare the last two rows first: weighting by normalized probability gains nothing over simply counting. The paper's explanation is that the model rates all its sampled paths as similarly likely, right and wrong alike: it is poorly calibrated on this task, so its confidence adds no information. Unnormalized weights do much worse; a likely reason is that products shrink with every token, so long paths are drowned out. Averaging instead of summing is worse still, because it discards agreement, as the tiny example showed.
Why it matters
If the model cannot tell its right paths from its wrong ones, something else has to: agreement between paths (this paper), or a separately trained verifier that scores each path, as the paper notes earlier work did. The reasoning lesson's verifier_experiment compares the two.
Try it: when does voting help?
Everyday picture
A vote picks the biggest pile, not the right answer. It finds the right answer only when the right answer's pile tends to be bigger than every wrong answer's pile. That depends on two things: how often one path is right, and how scattered the wrong paths are.
Tiny example
A model right 40% of the time on a yes/no question has only one wrong answer, taking the other 60%: voting makes it worse. The same 40% model on a maths question, with its mistakes spread over ten different wrong numbers (6% each), sees the right answer as the biggest pile, and voting lifts it towards 100%.
Try it: drag the sliders. Watch the solid line (the vote) against the dashed line (one sampled path). Then set the wrong answers to 1 and move the first slider below and above 50%.
Hover the chart, or tab to it and use the arrow keys, to read the accuracy at each number of paths.
Reading it: the x-axis is m, the number of sampled paths that vote, from 1 to 40; the y-axis is how often the vote returns the right answer. The dashed line is a single sample, flat at the first slider's value. At the start (40% right, mistakes over ten answers), each wrong answer gets only 6%, so the right answer is the biggest pile and the vote climbs past 90% by about fifteen paths, then flattens, the same shape as the paper's Figure 2. Set the wrong answers to 1 (a yes/no question) and the vote rises only above 50%; below it, the vote sinks. The rule behind every setting: voting helps exactly when the right answer is more likely than each single wrong answer, p > (1 − p) / k. At p = 0.25 with k = 3, all four answers are equally likely and the line stays flat at 0.25. The curves are computed exactly, assuming paths are independent; correlated_vote_accuracy in the reasoning lesson shows what shared mistakes do to them.
Why it matters
This is why self-consistency shines on maths, where wrong answers are many and scattered, and helps less on yes/no and multiple-choice tasks, where a common wrong answer can collect as many votes as the right one. The paper's own commonsense gains (ARC-c +3.9) are smaller than its arithmetic ones (GSM8K +17.9).
3 Experiments · original
Everyday picture
The test is fair by construction: the same models, the same prompts and the same questions as chain-of-thought prompting. The only change is how the output is decoded: one greedy chain, or 40 samples and a vote.
3.1 Experiment setup · original
Everyday picture
Four models of different sizes sit the same papers: arithmetic word problems, commonsense questions and two symbol puzzles.
What was tested
- Tasks. Arithmetic: AddSub, MultiArith, ASDiv, AQuA, GSM8K, SVAMP. Commonsense: CommonsenseQA, StrategyQA, ARC (easy and challenge). Symbolic: last-letter concatenation and coin flip, tested at 4 letters and 4 flips when the prompt shows only 2.
- Models. UL2 (20B, open), GPT-3 through the Codex engines code-davinci-001 and code-davinci-002, LaMDA-137B and PaLM-540B.
- Prompts. The chain-of-thought paper's own: the same 8 worked examples for every arithmetic task; 4 to 7 per commonsense task.
- Sampling. Temperature T = 0.5 with top-k k = 40 for UL2 and LaMDA; T = 0.7, k = 40 for PaLM; T = 0.7 without top-k for GPT-3.
- Protocol. 40 sampled outputs per question, and every result averaged over 10 runs.
Why it matters
These are ordinary settings for open-ended text, not tuned for reasoning. Section 3.5 shows the results hold across temperatures and sampling schemes, so the gain is not an artefact of one lucky setting.
3.2 Main results · original
Everyday picture
A good student gains a little from double-checking; a middling one gains a lot. But a student who cannot yet do the sums gains little either, because their attempts don't agree on anything.
Tiny example
On GSM8K, PaLM-540B goes from 56.5% with one greedy chain to 74.4% with the vote: +17.9 points. UL2-20B goes from 4.1% to 7.3%: when almost every path is wrong, there is little right answer to amplify.
Reading it: each pair of bars is one arithmetic benchmark, axis 0 to 100%: plain grey for chain of thought with one greedy chain, striped blue for self-consistency with 40 sampled chains. Every striped bar is longer than its grey partner, for every model: the vote never lost. Switch models and compare the gaps: small for UL2-20B (+3 to +7), large for LaMDA-137B (+9 to +24), and still large on the hardest sets for the strongest models (PaLM-540B: +12.5 on AQuA, +17.9 on GSM8K). Numbers from Table 2 of Wang et al. (2022), reproduced under CC BY 4.0. (The chain-of-thought paper reports PaLM-540B at 56.9% on GSM8K; this paper's own greedy run gives 56.5%.)
Commonsense and symbolic tasks
Gains are smaller, but never negative (Table 3): the smallest are zero, on 4-letter concatenation for UL2-20B and LaMDA-137B, which can barely do the task at all. With code-davinci-002: StrategyQA 73.4% → 79.8%, ARC-challenge 83.6% → 87.5%. On the longer symbol puzzles, PaLM-540B gains 5.0 points on 4-letter concatenation and 3.0 on 4-flip coins.
How many paths?
The paper plots accuracy against 1, 5, 10, 20 and 40 sampled paths (its Figure 2). The curves rise steeply over the first few paths and then level off, much like the vote explorer above. One detail matters: a single sampled path is usually a little worse than the greedy one, because sampling adds noise. The vote is what turns that noise into an advantage.
A repaired example
| Decode | Output |
|---|---|
| Greedy | He traveled 60 miles in total. He stopped after 20 miles, so he traveled 60 - 20 = 40 miles between the first and second stops. The answer is 40. |
| Sampled path 1 | … So between his first and second stops he travelled 60 - 20 -15 = 25 miles. The answer is 25. |
| Sampled path 2 | … This means the second stop was 60 - 15 = 45 miles into the trip. Since he made the stops in order, the second stop must have been 45 - 20 = 25 miles after the first stop. The answer is 25. |
Why it matters
Greedy decoding commits to the most likely next token at every step, which is not the same as the most likely answer. Here the greedy path forgets the second stop; two differently worded samples both remember it and agree on 25.
3.3 Self-consistency helps when chain of thought hurts · original
Everyday picture
Asking someone to explain every answer can make them overthink a question they would have got right on instinct. Asking them several times and taking the consensus smooths that out.
Tiny example
On RTE, a natural language inference task, PaLM-540B scores 84.8% with plain answers, 79.1% with one chain of thought, and 86.3% with self-consistency.
| Method | ANLI R1 | e-SNLI | RTE | BoolQ |
|---|---|---|---|---|
| Standard prompting (no chain) | 69.1 | 85.8 | 84.8 | 71.3 |
| Chain-of-thought prompting | 68.8 | 81.0 | 79.1 | 74.2 |
| Self-consistency | 78.5 | 88.4 | 86.3 | 78.4 |
Reading it: in the first three columns the middle row is below the top row: adding a single chain of thought hurt, as Ye and Durrett (2022) had reported. The bottom row beats both on every column. So the paper recommends self-consistency as the safe way to add reasoning to a few-shot prompt.
Why it matters
One chain is one draw from a noisy process. Whether reasoning helps or hurts on average can be hidden by that noise; the vote averages it out.
3.4 Compared with other approaches · original
Everyday picture
There are other ways to get more than one attempt. You could keep the attempt the model is most sure of; explore several continuations at once and keep the best; or ask with differently worded prompts. The paper tries each against the vote, with the same number of attempts.
Tiny example: three rivals
- Sample-and-rank (code-davinci-001): sample the same number of outputs, return the one with the highest probability. It helps a little; the vote helps much more (the paper's Figure 3).
- Beam search (UL2-20B): keeps the few most likely partial outputs at each step. Its outputs are too alike to vote usefully, and its top beam gets worse as beams are added.
- Ensembles of prompts (LaMDA-137B): 40 random orderings of the exemplars, or 3 hand-written prompt sets, each decoded greedily, then a vote.
| Method | 1 | 5 | 10 | 40 |
|---|---|---|---|---|
| UL2-20B on AQuA, by number of beams or paths | ||||
| Beam search (top beam) | 23.6 | 19.3 | 16.1 | 10.2 |
| Self-consistency over beam-search outputs | 23.6 | 19.8 | 21.2 | 24.2 |
| Self-consistency over samples | 19.7 | 24.9 | 25.3 | 26.9 |
| LaMDA-137B | GSM8K | MultiArith | ARC-c |
|---|---|---|---|
| Chain of thought, greedy | 17.1 | 51.8 | 55.1 |
| Ensemble of 3 prompt sets | 18.6 | 57.1 | 57.0 |
| Ensemble of 40 prompt orderings | 19.2 | 60.9 | 57.0 |
| Self-consistency, 40 paths | 27.7 | 75.7 | 59.8 |
Reading it: in the first table, read along each row. The top beam falls from 23.6% to 10.2% as the beam widens; voting over beams dips, then recovers to about where it started; voting over samples starts lower (one sample is noisier than the top beam) and climbs to the best score. In the second table, the prompt ensembles gain 1 to 9 points over greedy chain of thought, while self-consistency gains 5 to 24. The lesson in both: what makes the vote work is diversity of reasoning paths, and sampling from one prompt gives more of it than reshuffling prompts or searching for the single best continuation.
Why it matters
The paper calls self-consistency a “self-ensemble”. Its appendix also tried ensembling three different models by voting over their greedy answers: 33.3% on GSM8K, far below PaLM-540B's own 74.4% with self-consistency, because weaker models drag the vote down.
3.5 Additional studies · original
Everyday picture
A method you have to tune carefully is fragile. The paper varies the sampling settings, deliberately corrupts the prompt, and asks whether agreement between samples can serve as a confidence meter.
Robust to sampling settings and model scale
Across temperatures, top-k values and nucleus sampling settings on PaLM-540B, and across LaMDA models of every size, self-consistency improved over greedy decoding (its Figure 4). Gains are smaller for small models, which cannot yet do the arithmetic.
Robust to imperfect prompts, and to zero-shot
| Model and prompt | Greedy | + Self-consistency (40) |
|---|---|---|
| LaMDA-137B, correct chains in the exemplars | 17.1 | 27.7 |
| LaMDA-137B, exemplars with wrong numbers in their chains | 14.9 | 23.4 |
| LaMDA-137B, exemplars showing equations only | 5.0 | 6.5 |
| PaLM-540B, zero-shot “let's think step by step” | 43.0 | 69.2 |
Reading it: the second row replaced every number in the exemplars' chains with a random one (“There are 7 cars in the parking lot already. 6 more arrive. Now there are 7 + 6 = 5 cars.”) but kept the final answers. Greedy accuracy dropped, and the vote more than recovered it. Equations-only chains are short, so samples differ little and the vote gains little. The last row needs no worked examples at all, only the instruction to think step by step (Kojima et al., 2022), and gains 26.2 points. The first row's 27.7 is Table 2's 40-path result, repeated for comparison.
The math: consistency as a confidence meter
In words: “the share of the sampled paths that agree with the winning answer.”
With the numbers: Figure 1's vote is 2 of 3, so consistency is 2 / 3 = 0.67. Forty paths with 34 on the winner give 0.85; with 12 on the winner, 0.30.
In Python:
answers = [18, 26, 18]
m = len(answers)
# max_a Σ_i 1(a_i = a): the winner's vote count
top = max(sum(1 for a_i in answers if a_i == a) for a in set(answers))
round(top / m, 2) # → 0.67
# forty paths, 34 or 12 of them on the winner
34 / 40, 12 / 40 # → (0.85, 0.3)
On GSM8K the paper found this share tracks accuracy closely (its Figure 5): questions where most paths agree are usually answered correctly, and questions with scattered answers usually are not. So low consistency is a cheap signal that the model “doesn't know”.
Why it matters
This is the useful by-product: a confidence score that needs no access to the model's internals, only its answers. In practice, low agreement is a natural trigger to escalate: sample more, call a tool, use a stronger model, or ask a person. The ReAct companion shows exactly that: its authors fall back from self-consistency to looking things up when fewer than half the samples agree.
4 Related work · original
The paper sets itself against three kinds of earlier work. Specialised systems for arithmetic or logic needed task-specific designs. Re-ranking methods trained an extra model to pick among samples: a verifier for maths solutions (Cobbe et al., 2021) or a human-labelled re-ranker for dialogue (LaMDA). Methods for extracting reasoning paths needed task-specific training too. Self-consistency needs none of that: sampling and counting on top of one frozen model. It also uses “consistency” in a new sense: agreement between the final answers of different reasoning paths.
5 Conclusion and discussion · original
“One limitation of self-consistency is that it incurs more computation cost.”Wang et al. (2022), §5
Everyday picture
Asking forty people costs forty times as much as asking one. The paper's practical advice: start with 5 or 10 paths, which capture most of the gain, since accuracy levels off quickly.
The math: what the vote costs
In words: “the tokens generated for one question are the number of sampled paths times the length of each.”
With the numbers: the paper capped GPT-3 outputs at 128 tokens. One greedy chain: at most 128 tokens. 40 paths: at most 40 × 128 = 5,120. The suggested 5 or 10 paths: 640 or 1,280. Run side by side, the wait is still about one path long; the bill is m paths.
In Python:
# T: the paper's cap on GPT-3 output length
T = 128
# m = 40 paths, as in the experiments
40 * T # → 5120
# the suggested starting points
[m * T for m in (5, 10)] # → [640, 1280]
Other points in the discussion
- Sampled chains that reach the right answer could become training data, so a model learns to get it right in one pass.
- Right answers can still ride on wrong reasoning: in Table 4's StrategyQA example, the populations the model quotes for the two Albanys are not accurate, though its answer is.
- Agreement between samples gives an uncertainty estimate and better calibration, for free.
Why it matters today
The first bullet came true: training on a model's own correct sampled chains, and later reinforcement learning that rewards right final answers, turned sampling-and-checking into models that reason well in one pass. The cost trade-off, money for accuracy, is the everyday arithmetic of test-time compute; the reasoning lesson prices it (reasoning_cost).
What changed since 2022
| In the paper | Today | Learn it |
|---|---|---|
| Vote over final answers | Still the default when answers compare exactly (maths, multiple choice, code outputs), often called majority voting | reasoning |
| Agreement picks the answer | A verifier picks it (best-of-n), bounded above by pass@n | pass@n |
| Independent samples assumed | Known limit: samples from one model share its blind spots, so gains level off | correlated votes |
| Consistency as a confidence meter | Agreement used to decide when to escalate, retry or ask a person | ReAct companion |
| 40 paths from a prompted model | Models trained to reason, which check their own work inside one long chain | chain-of-thought companion |
Glossary
Every term with hover guidance on this page, in one place.