Part of Language AI Handbook
Explains how chain-of-thought prompting enables language models to reason step by step. Topics include few-shot CoT, zero-shot CoT, self-consistency.
Choose your expertise level to adjust how many terms are explained. Beginners see more tooltips, experts see fewer to maintain reading flow. Hover over underlined terms for instant definitions.
Article links
Make inline references clickable
Chain-of-Thought Prompting
When you ask a language model a hard question, it has two choices. It can jump straight to an answer, pattern-matching against whatever surface features most resemble the question. Or it can work through the problem step by step, articulating intermediate reasoning before committing to a conclusion. Chain-of-thought (CoT) prompting is a technique that coaxes models into the second behavior. By showing a model examples where reasoning is explicit and sequential, or simply by asking it to "think step by step," you can dramatically improve performance on tasks that require multi-step logic.
The core insight is deceptively simple: language models generate text token by token, and the tokens they produce become context for subsequent tokens. When a model writes out an intermediate reasoning step, that step is available as context when it generates the next step. The scratchpad directly shapes what the model can compute.
Chain-of-thought prompting was introduced by Wei et al. (2022) in a landmark paper demonstrating that large models given reasoning-rich few-shot examples substantially outperformed those given only input-output pairs, particularly on arithmetic, commonsense, and symbolic reasoning tasks. The improvement was striking enough to reframe how researchers thought about in-context learning: the format of examples matters as much as the examples themselves.
Before CoT, scaling up models was the dominant strategy for improving reasoning performance. You could throw a harder question at a bigger model, and bigger models tended to do better. CoT changed the equation by showing that how you prompt matters just as much as how large your model is. A well-designed CoT prompt on a 100-billion-parameter model can outperform a poorly designed prompt on the same model by a wider margin than simply doubling the model size would produce. That realization transformed prompt engineering from a folklore art into an active research discipline.
The Historical Context Behind Chain-of-Thought
To appreciate why CoT was such a breakthrough, it helps to understand what came before it. Prior to 2022, the prevailing approach to few-shot in-context learning was simple: give the model a small number of question-answer pairs and hope it learned the task from those examples. For factual retrieval, classification, and simple pattern-matching tasks, this worked reasonably well. For multi-step reasoning, it failed almost uniformly.
Researchers had tried various approaches to patch this gap. Some added a scratchpad area to the prompt, allowing the model to write intermediate work before committing to an answer. Others used output-parsing strategies that broke down a single complex question into multiple simpler subquestions. Still others relied on careful problem decomposition by humans before passing anything to the model. Each approach had merit, but none combined the simplicity, generality, and effectiveness that Wei et al. would achieve.
The Wei et al. (2022) paper made several specific observations that helped explain the gap. First, the improvement from CoT was not uniform across model sizes. Small models did not benefit, and in some cases degraded with CoT. The benefit emerged sharply around 100 billion parameters and grew with scale. This emergent property suggested that CoT was eliciting a latent capability that only manifested in sufficiently large models. Second, the improvement was especially pronounced for tasks with a natural step-by-step structure: arithmetic word problems, commonsense causality questions, and symbolic manipulation tasks. For tasks that did not require sequential reasoning, such as extractive question answering, the benefit was smaller.
Simultaneously, Kojima et al. (2022) published a companion paper showing that you did not need carefully crafted exemplars at all. The single phrase "Let's think step by step," appended to the question before the model's response, was enough to elicit coherent reasoning chains. This zero-shot finding was arguably even more surprising than the few-shot result. It demonstrated that large language models already possessed latent reasoning capabilities and that eliciting them required only a minimal cue.
Together, these two papers marked a watershed moment. Reasoning capability was no longer an innate property of a model that either existed or did not. It was a behavior that prompt designers could reliably elicit, shape, and improve.
Standard Prompting vs. Chain-of-Thought
To understand what CoT adds, it helps to see standard prompting and CoT prompting side by side on the same problem.
Standard few-shot prompting shows the model input-output pairs:
Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls.
Each can has 3 balls. How many tennis balls does he have now?
A: 11
Q: The cafeteria had 23 apples. If they used 20 to make lunch and
bought 6 more, how many apples do they have?
A:
The model sees that the answer to the first question is 11 and learns to pattern-match toward a number. But it has no explicit record of how that number was derived. When the model generates an answer to the cafeteria question, it is working purely from statistical patterns: questions like this tend to have small integer answers, and the numbers in the question provide anchors. This works well when the required computation is simple and fits within a single inference step. For more complex problems, the failure mode is that the model produces a plausible-sounding number that is statistically consistent with the question format but arithmetically wrong.
Chain-of-thought prompting adds reasoning steps to each example:
Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls.
Each can has 3 balls. How many tennis balls does he have now?
A: Roger started with 5 balls. 2 cans × 3 balls = 6 balls.
5 + 6 = 11. The answer is 11.
Q: The cafeteria had 23 apples. If they used 20 to make lunch and
bought 6 more, how many apples do they have?
A:
Now the model sees that good answers include explicit arithmetic. When it generates an answer, it is likely to also generate reasoning steps, and those steps help it arrive at the correct final number. For the cafeteria problem, a CoT-conditioned model will typically produce "The cafeteria started with 23 apples. They used 20, leaving 23 - 20 = 3 apples. Then they bought 6 more: 3 + 6 = 9 apples." The reasoning is explicit, verifiable, and structured.
The performance gap between these two formats grows with problem difficulty. On simple one-step arithmetic, both approaches work comparably. On multi-step word problems or questions requiring multiple inference steps, CoT substantially outperforms standard prompting, and the advantage widens as problem complexity increases. Wei et al. reported accuracy improvements of over 40 percentage points on the GSM8K grade-school math benchmark for large models, a gap that no amount of prompt engineering on the answer-only format could close.
A less obvious benefit of the explicit reasoning format is that wrong answers become diagnosable. With standard prompting, you see "17" when the answer is 9, and you have no idea whether the model misread the problem, performed incorrect arithmetic, or simply pattern-matched to a wrong number. With CoT, you see the step where things went wrong. This turns debugging from a black-box exercise into something resembling reading a student's test paper.
Why Chain-of-Thought Works
The improvement from CoT is not just behavioral. It reflects something real about how autoregressive language models compute, and understanding the mechanism helps you apply and extend CoT more effectively.
The Scratchpad as Extended Computation
Transformers process input in parallel within each forward pass, but they generate output sequentially, one token at a time. Each generated token extends the context window and becomes part of the input for the next token. This means the model's effective computation per output token is bounded by how much the attention mechanism can integrate within a single forward pass.
When intermediate reasoning steps are written out explicitly, the model does not need to compress all necessary computation into a single prediction. Instead, it distributes the work across multiple generation steps. Each step is a new forward pass with more context, including the reasoning produced so far.
Consider a problem that requires four sequential arithmetic operations. Without CoT, the model must implicitly perform all four in the course of predicting one answer token. The residual stream at that final token must somehow carry the result of four cascaded computations, each dependent on the previous. This places severe demands on the model's internal representations. With CoT, the first operation is performed and written out, then the second operation can take the written result as input at the token level, at both the token and representation levels. The problem is decomposed into four separate computation steps.
This decomposition has formal parallels to circuit complexity theory. A problem that requires sequential computations, where each computation depends on the result of the previous one, has a serial depth of . A model with fixed depth (say, transformer layers) can handle computations of depth up to in a single forward pass. For problems with depth exceeding , the model cannot perform the computation in one pass. CoT effectively increases the usable computation depth by breaking the problem across multiple passes: each generation step is a fresh -layer computation over the context that includes all previous reasoning. The effective depth is , where is the number of reasoning steps. This analysis, formalized by Feng et al. (2023) among others, provides a theoretical basis for why CoT dramatically improves performance on problems that require many sequential steps.
In-Context Learning and Reasoning Format
From the in-context learning perspective, CoT prompting teaches the model which program to run by example. Standard few-shot examples teach the model: "respond to a question with a short answer." CoT examples teach it: "respond to a question with a reasoning chain followed by an answer."
Language models trained on large text corpora have seen millions of examples of human reasoning, from math textbook solutions to forum explanations to scientific papers. CoT prompting activates this latent capability by demonstrating the desired output format. The model generalizes from the few-shot examples to apply the same reasoning pattern to new questions.
This framing helps explain why the improvement is not uniform across model sizes. Smaller models have not seen enough training data, have insufficient capacity to reliably identify and apply reasoning formats from few-shot examples, or lack the representations necessary to perform multi-step inference even when prompted. The capability is absent rather than hidden behind a prompt barrier. At sufficiently large scale, the capability exists and the prompt elicits it. This is why CoT's effectiveness is considered an emergent property: it appears abruptly and qualitatively differently above a scale threshold rather than improving gradually.
Error Localization and Correction
Explicit reasoning chains have a practical benefit that pure answer generation does not: they are interpretable. When a model produces an incorrect answer in standard prompting, it is hard to know where the reasoning went wrong, or even that there was reasoning to inspect. With CoT, you can read the chain and identify exactly where the model made an error, whether in problem decomposition, arithmetic, logical inference, or fact retrieval.
This interpretability also matters during development. Researchers can examine failing CoT traces to understand failure modes, identify categories of errors, and design better prompting strategies or training procedures accordingly. If you see that a model consistently fails on CoT chains that require percentage calculations, you know to add exemplars that demonstrate percentage arithmetic. If you see errors only in chains where a variable's value must be tracked across four or more steps, you know to add an explicit bookkeeping step to your exemplars.
Error localization also enables human-in-the-loop correction. A human reviewer can spot the error in a CoT chain, correct it, and either use the corrected chain directly or incorporate it as a new exemplar. This creates a feedback loop that is much more tractable than trying to correct output by adjusting prompts blindly.
Few-Shot Chain-of-Thought
The original CoT formulation is a few-shot approach. You provide between four and eight exemplar question-answer pairs, each with an explicit reasoning chain, before the question you want answered. The model learns from these exemplars to produce similar reasoning chains when generating its own answer.
The choice of four to eight exemplars reflects a practical balance. Too few exemplars, and the model may not clearly understand the desired reasoning format. Too many, and the prompt grows unwieldy and begins to eat into the space available for the target question and its answer. For most tasks, four exemplars reliably establish the format, and diminishing returns kick in beyond eight.
Exemplar Construction
The quality of CoT exemplars significantly affects performance. Several principles guide their construction.
The reasoning chains should be correct. Incorrect exemplars teach the model wrong reasoning patterns and often degrade performance below even standard prompting. This seems obvious but matters practically: when building exemplars by hand, small arithmetic errors or logical slips are easy to introduce. Always verify every step in every exemplar numerically, not just by inspection.
The reasoning should be atomic. Each step should represent one conceptual operation: a single arithmetic calculation, one logical deduction, one fact retrieval. Long compound steps that secretly perform multiple operations at once reduce the benefit of writing steps out, since the model still has to perform all the work in one generation step to produce that compound step. A step like "She earns $12 per hour, working 8 hours per day for 5 days over 3 weeks, so she earns 12 × 8 × 5 × 3 = $1440" defeats the purpose. Better to write "She earns 12 × 8 = $96 per day. Over 5 days: 96 × 5 = $480 per week. Over 3 weeks: 480 × 3 = $1440."
The exemplars should cover the reasoning patterns expected in the test set. If the test questions involve percentage calculations, some exemplars should demonstrate percentage reasoning. If the test questions involve temporal ordering, some exemplars should demonstrate that. Diverse exemplars generalize better than redundant ones that all demonstrate the same pattern. If all four of your exemplars involve simple addition and subtraction, the model will not learn to apply CoT to questions requiring multiplication or division.
Exemplar annotation is expensive. For a complex domain, constructing eight exemplars with high-quality reasoning chains requires significant expert effort. This limitation motivated research into reducing the annotation burden, ultimately leading to zero-shot CoT and automated exemplar generation.
Selecting Exemplars Automatically
Not all exemplars are equally useful. Automatic exemplar selection methods choose questions most likely to activate the relevant reasoning skills. Active prompting (Diao et al., 2023) identifies the questions where a model is most uncertain, measured by variance across multiple sampled answers, and solicits human annotations for those. The intuition is that these uncertainty points correspond to reasoning patterns the model needs the most help with. You are using the model's own uncertainty to guide where human annotation effort goes.
Complexity-based selection (Fu et al., 2023) prefers exemplars with longer, more detailed reasoning chains, on the theory that more complex exemplars teach more thorough reasoning. Empirically, selecting for chain complexity often outperforms random selection. When given a pool of candidate exemplars varying in step count from two to eight, selecting the four with the longest chains typically outperforms randomly selecting four. The model learns to be more thorough when it sees thorough exemplars.
Diverse selection, which explicitly picks exemplars covering different reasoning patterns and problem types, consistently outperforms naive selection. If you have a budget of four exemplars for a math word problem task, choosing one that demonstrates rate calculations, one that requires unit conversion, one involving sequential accumulation, and one involving distribution tends to outperform choosing four problems all of the same type.
The Role of Exemplar Order
The order in which exemplars appear in the prompt has a modest but measurable effect on performance. Models tend to be somewhat recency-biased, meaning the last exemplar can have slightly more influence on the output format. For most tasks, the effect is small enough to ignore, but when fine-tuning prompts for critical applications, experimenting with exemplar order is worthwhile. Placing the most relevant or most complex exemplar last is a reasonable heuristic.
Zero-Shot Chain-of-Thought
Few-shot CoT requires constructing exemplars for each new domain or task. Zero-shot CoT, introduced by Kojima et al. (2022), eliminates this requirement entirely with a single addition to the prompt: the phrase "Let's think step by step."
Appending this phrase before the model's response elicits reasoning chains without any demonstration examples. The model has encountered this phrase in training data in contexts where it introduces structured reasoning, so it generalizes to produce similar reasoning structures in new contexts. The mechanism is essentially retrieval: by pattern-matching to contexts in pretraining data where careful step-by-step explanation follows, the model activates the register of careful explanation rather than the register of quick-answer generation.
The Two-Stage Pipeline
Zero-shot CoT is implemented as a two-stage process. In the first stage, the reasoning prompt asks the model to think through the problem:
Q: If you have 3 liters of water and drink half of it, then add 0.5
liters, how much water do you have?
A: Let's think step by step.
The model generates a reasoning chain in response. In the second stage, a separate extraction prompt retrieves the final answer:
[original question]
[generated reasoning chain]
Therefore, the answer (arabic numerals) is
The two-stage design matters because language models asked to reason and answer simultaneously sometimes produce reasoning that drifts from the eventual answer. By separating reasoning from answer extraction, the pipeline ensures the final answer is grounded in the completed chain rather than generated in parallel with it. The extraction step forces the model to commit to a specific answer based on what the reasoning concluded, rather than interpolating between the reasoning output and whatever the model would have produced without reasoning.
In practice, when you call an LLM API, you implement this as two sequential API calls. The first call returns a reasoning chain. You concatenate the original question, the reasoning chain, and a prompt like "Therefore, the answer is" and send that as the second call. The second call's first generated tokens constitute the answer.
Why "Let's Think Step by Step" Works
The phrase "Let's think step by step" appears in training data in a specific kind of context: tutorial explanations, worked examples, educational content, collaborative problem-solving sessions, and instructional writing of all kinds. By activating these contexts, it shifts the model toward a reasoning-oriented generation mode, one associated with carefulness, thoroughness, and structured exposition.
Kojima et al. (2022) tested numerous other trigger phrases and found significant variation in effectiveness. "Let's think step by step" consistently performed best across tasks, though phrases like "Let's work this out in a few steps" and "Let's think about this carefully" also improved performance over no trigger. The specific phrasing matters because different phrases activate different distributions in the model's training data. "Let's think step by step" is particularly common in educational and instructional writing, which tends to contain exactly the kind of careful, sequential reasoning that helps with multi-step problems. "Think about this question" activates a broader distribution that includes casual conversation and opinion-sharing, which is less helpful.
The gap between zero-shot and few-shot CoT reveals an important asymmetry. Zero-shot CoT gives the model full latitude to decide what "step by step" means for a given problem. Few-shot CoT constrains the format through demonstration, which is helpful when the appropriate reasoning pattern is specific and non-obvious, but adds overhead. For well-structured mathematical problems, few-shot CoT with good exemplars consistently outperforms zero-shot CoT. For open-ended reasoning tasks where the relevant steps vary unpredictably, zero-shot CoT can match or exceed few-shot CoT by allowing the model to adapt its reasoning format to the specific question.
Zero-shot CoT is remarkably effective given that it adds just a few words to the prompt. On the GSM8K grade school math benchmark, it improved GPT-3 performance from around 10% to around 40% accuracy. It performs somewhat below well-constructed few-shot CoT on the same benchmarks, but the trade-off in annotation effort is dramatic: zero annotation versus expert construction of 8 high-quality exemplars.
Self-Consistency
A natural extension of chain-of-thought prompting is to generate not one reasoning chain, but many, and then aggregate their conclusions. Self-consistency (Wang et al., 2022) does exactly this. The model samples multiple diverse reasoning chains for the same question, and the final answer is determined by majority vote over the conclusions of those chains.
The intuition is that there are many valid ways to reason toward a correct answer, but fewer ways to reason toward any specific incorrect answer. Errors in reasoning chains tend to be idiosyncratic, arising from specific wrong assumptions, arithmetic mistakes, or misreadings that differ from chain to chain. The correct answer, by contrast, tends to be reached by multiple independent chains that decompose the problem differently but converge on the same conclusion. Majority vote over conclusions therefore favors the correct answer.
This framing makes a probabilistic prediction: the correct answer should be the mode of the distribution over sampled answers, and this mode should become more reliable as the number of samples increases. Empirically, this prediction holds well. The improvement from self-consistency is consistent across model sizes, tasks, and benchmarks, and continues to grow with sample count, though with diminishing returns.
Self-consistency substantially improves accuracy over single-chain CoT, particularly on mathematical and symbolic reasoning tasks. On GSM8K, self-consistency with 40 sampled chains improves accuracy by 10-20 percentage points over single-chain CoT with the same model. The improvement is most pronounced for problems where CoT single-chain accuracy is in the 50-70% range, exactly the range where errors are common but not universal.
Temperature and Diversity
For self-consistency to work, the sampled chains must be diverse. If you sample with temperature 0 (greedy decoding), every chain is identical and you gain nothing. You need to use a nonzero temperature to introduce variation across samples. Wang et al. (2022) found that temperatures between 0.5 and 0.8 work well: high enough to produce distinct reasoning paths but not so high that chains become incoherent.
The diversity of chains matters for correctness voting and for the richness of reasoning strategies explored. At a moderate temperature, different sampled chains may approach the same problem via different decompositions, catching errors that any single decomposition might miss. One chain might compute earnings per day first, then per week, then per three weeks. Another might compute total hours worked first, then apply the hourly rate. Both arrive at the same answer via different arithmetic paths, and their agreement provides stronger evidence of correctness.
Computational Trade-offs
The cost is computational. Sampling 40 reasoning chains multiplies inference cost by 40. This trade-off is worthwhile when accuracy is critical and latency is flexible, but impractical for real-time applications. Most production deployments that use self-consistency settle on 8-16 samples, capturing most of the accuracy improvement while keeping costs manageable. The self-consistency accuracy curve shows strong diminishing returns beyond 16 samples, so few applications justify the full 40-sample budget.
Universal self-consistency (Chen et al., 2023) proposed a more efficient variant: rather than running full generation for every chain, use the first-pass reasoning chains to identify low-confidence questions (those with high variance in sampled answers) and apply multi-sampling only to those. Questions with high first-pass confidence do not benefit much from additional samples, so the budget can be concentrated where it helps most.
Least-to-Most Prompting
One limitation of standard CoT is that it generates a single linear chain that handles the full problem in one pass. For questions involving explicit problem decomposition, a more structured approach can help. Least-to-most prompting (Zhou et al., 2022) breaks the task into two stages: first decompose the problem into subproblems, then solve them sequentially, using previous solutions as context for subsequent ones.
The first stage uses a decomposition prompt to identify the constituent subproblems:
Q: How long does it take to travel 150 miles if you drive at 60 mph
for the first half and 50 mph for the second half?
To solve this problem, we first need to know:
The model completes this with subquestions like "What distance is covered at each speed?" and "How long does each segment take?" The second stage then solves each subquestion in order, with the answer to each available as context when addressing the next.
This structured decomposition is especially useful for problems where the subproblems have a clear dependency structure. Least-to-most prompting generalizes better to longer and more complex versions of problems than single-chain CoT. Where standard CoT might succeed on a two-step problem and fail on a five-step version, least-to-most can handle the five-step version by explicitly tracking the dependency structure. The model never has to hold five simultaneous threads in a single inference step: it proceeds one thread at a time, with the result of each step written into the context before the next is attempted.
Zhou et al. (2022) showed strong generalization results on the SCAN compositional generalization benchmark. Models using least-to-most prompting generalized to instructions with significantly more steps than those seen in training examples, while standard CoT failed to generalize. The key difference is that least-to-most makes the compositional structure explicit, allowing the model to use individual subproblem solutions rather than trying to handle the entire compositional challenge at once.
Program-Aided Reasoning
A significant limitation of all purely textual CoT approaches is that they rely on the language model to perform arithmetic and symbolic manipulation reliably. For arithmetic especially, this reliability is limited: even large models make arithmetic errors in multi-step calculations, particularly when numbers are large or intermediate results must be tracked carefully.
Program-aided language models (PAL, Gao et al., 2022) address this limitation by delegating the computational steps to an interpreter. Rather than generating a natural language reasoning chain and extracting a number at the end, the model generates Python code that performs the reasoning, and the code is then executed to produce the final answer.
A PAL response to the muffin problem might look like:
# A bakery makes 48 muffins per batch
muffins_per_batch = 48
# 3 morning batches and 2 afternoon batches
morning_batches = 3
afternoon_batches = 2
total_batches = morning_batches + afternoon_batches
# Total muffins produced
total_muffins = muffins_per_batch * total_batches
# Sell 85
sold = 85
# Remaining
remaining = total_muffins - sold
print(remaining)155
The Python interpreter executes this and prints the answer. The language model's job shifts from performing arithmetic to writing code that correctly models the problem structure. This plays to the model's strengths: large models are excellent at code generation and considerably less reliable at mental arithmetic for large numbers.
PAL achieves substantially higher accuracy on arithmetic benchmarks than standard CoT, with improvements of 10-20% on GSM8K depending on the base model. The approach is particularly effective for problems involving large numbers, repeated operations, or tracking multiple variables, exactly the cases where CoT arithmetic errors accumulate.
The practical limitation is that PAL requires a Python interpreter in the inference pipeline and the ability to execute untrusted code safely. For production deployments, sandboxing this execution adds engineering complexity. The approach also assumes that the problem can be naturally expressed in procedural code, which holds for arithmetic and data-manipulation tasks but not for all reasoning domains.
Tree-of-Thought Prompting
Both CoT and least-to-most prompting generate linear sequences of reasoning steps. Tree-of-thought (ToT) prompting (Yao et al., 2023) generalizes this to a tree structure, where the model explicitly explores multiple reasoning paths, evaluates partial solutions, and prunes unpromising branches before committing to a final answer.
In ToT, the reasoning process involves three components. First, a thought generator proposes multiple candidate next steps (thoughts) from the current state. These might be different problem decompositions, alternative arithmetic approaches, or distinct subproblem orderings. Second, a state evaluator assesses the quality or promise of each candidate thought, either by asking the model to score it directly ("Is this a good direction? Yes/No/Maybe") or by running a separate value estimation. Third, a search algorithm (breadth-first, depth-first, or beam search) navigates the tree by following promising branches and backtracking from dead ends.
The power of ToT is most evident on tasks that require non-obvious exploration. The Game of 24, where you must find arithmetic operations combining four numbers to reach 24, is a canonical example. Single-chain CoT rarely finds solutions because the task requires exploring many combinations before finding one that works. ToT succeeds by explicitly generating candidate combinations, evaluating whether each brings the calculation closer to 24, and backtracking when a branch reaches a dead end.
ToT comes at significant computational cost. Exploring a tree with branching factor three and depth four requires up to 81 model calls, compared to one for standard CoT. For most applications, this is prohibitive. But for high-stakes problems where accuracy matters more than cost, and for research into model reasoning capabilities, ToT provides a principled framework for extending CoT to problems requiring explicit search.
CoT Fine-Tuning
Few-shot and zero-shot CoT keep the model weights fixed and change only the prompt. An alternative is to train on reasoning chains directly, teaching the model to produce them as part of its standard output. CoT fine-tuning includes reasoning chains in the training data and optimizes the model to generate them.
Training on Reasoning Chains
Fine-tuning on CoT data was explored in the context of instruction tuning. Models fine-tuned on instruction datasets that include reasoning chains, such as chain-of-thought formatted examples from GSM8K, MATH, and other reasoning benchmarks, learn to spontaneously produce reasoning when asked questions that benefit from it.
The Flan series of models (Wei et al., 2022b; Chung et al., 2022) demonstrated that including chain-of-thought examples in instruction fine-tuning data improves reasoning performance even when reasoning is not explicitly requested at inference time. The model learns that careful, step-by-step reasoning is part of good response quality and applies it proactively. This is a meaningful shift: instead of requiring a trigger phrase or exemplars at inference time, the model develops a default behavior of reasoning thoroughly.
The benefit of CoT fine-tuning extends beyond the specific tasks included in training. Models trained on diverse reasoning chains from mathematics, science, and commonsense domains show improved reasoning on tasks that were not in the training mixture. The model appears to learn general reasoning habits, not just task-specific patterns. This transfer is more reliable when the fine-tuning data spans diverse reasoning types, and degrades when the training distribution is narrow.
Distillation from Large to Small Models
An important application of CoT fine-tuning is knowledge distillation. Large models generate high-quality reasoning chains for a training corpus, and a smaller model is trained to reproduce both the chains and the conclusions. This transfers some of the reasoning capability of the large model to the smaller one.
Ho et al. (2022) showed that fine-tuning smaller models on CoT rationales generated by a larger model substantially improves the smaller model's reasoning accuracy, often matching or exceeding what few-shot CoT with the larger model achieves. The fine-tuned small model also produces reasoning chains at inference time, which means its outputs remain interpretable.
The quality of the generated reasoning chains matters a great deal for distillation. Chains from a teacher model that sometimes err propagate those errors into the student. Filtering generated chains by correctness before fine-tuning, or using only chains that arrive at the correct final answer, consistently improves distilled model quality. A practical procedure: use the teacher model to generate 8-16 candidate chains per training example, keep only those that arrive at the verified correct answer, and use those verified chains as fine-tuning data. This filters noisy signal and leaves a cleaner, more reliable training corpus.
The theoretical picture behind distillation is interesting. Small models cannot reliably do multi-step reasoning via few-shot CoT because, as noted earlier, the capability is emergent and only manifests at large scale in the prompted setting. But fine-tuning changes the weights as well as the context. A small model trained specifically to produce reasoning chains can internalize the decomposition pattern as part of its weights, partially circumventing the scale threshold. The fine-tuned small model is not equivalent to a large model with CoT, but it substantially outperforms the fine-tuned small model without CoT, and often outperforms the large model without CoT.
Process Reward Models
A more principled approach to training on reasoning chains is to reward correct process, not just correct outcomes. Process reward models (PRMs), explored by Lightman et al. (2023), provide supervision at the step level. Human annotators label each step in a reasoning chain as correct or incorrect, and a reward model trained on these labels scores steps rather than only final answers.
This step-level supervision addresses a core problem with outcome-only training: a chain can arrive at the correct answer through flawed reasoning (getting lucky on the final step after errors earlier), or arrive at an incorrect answer despite mostly correct reasoning (a small error late in the chain). Outcome reward models reinforce both of these undesirable patterns. PRMs reward good reasoning regardless of whether the final answer happens to be correct, and penalize bad reasoning even when the answer is right.
The practical impact is most visible when PRMs are combined with search. Rather than sampling a single chain, the system uses the PRM to guide a search process such as beam search, Monte Carlo tree search, or best-of-N sampling that selects chains with high step-by-step quality. Each candidate step is scored by the PRM before the search continues, allowing the search to prefer branches that are both logically sound and factually accurate.
This produced state-of-the-art results on competitive math benchmarks as of 2024. OpenAI's work on process supervision (Lightman et al., 2023) showed that PRM-guided search on the MATH benchmark substantially outperformed both outcome-reward models and unguided self-consistency sampling. The improvement was especially pronounced on the hardest problems, where a single error in any step is fatal and the solution space requires broad exploration rather than reliable execution of a known procedure.
Implementation
Let's implement chain-of-thought prompting to see these ideas concretely. We'll build a few-shot CoT prompter, a zero-shot CoT two-stage pipeline, and a self-consistency aggregator, then compare them on a set of arithmetic word problems.
We'll define our test problems and exemplars. These exemplars are constructed with explicit, atomic reasoning steps:
# Arithmetic word problems for evaluation
test_problems = [
{
"question": "A bakery makes 48 muffins per batch. They run 3 batches in the morning and 2 in the afternoon. If they sell 85 muffins, how many are left?",
"answer": "155",
},
{
"question": "Sarah earns $12 per hour. She works 8 hours a day, 5 days a week. How much does she earn in 3 weeks?",
"answer": "1440",
},
{
"question": "A train travels 60 miles per hour. It departs at 9:00 AM and arrives at 1:00 PM. How far did it travel?",
"answer": "240",
},
{
"question": "A box holds 24 cans. A store receives 15 boxes and already had 36 cans in stock. They sell 120 cans. How many remain?",
"answer": "276",
},
{
"question": "A pool holds 2400 gallons. A pump fills it at 40 gallons per minute. How many minutes to fill the pool?",
"answer": "60",
},
]
# High-quality few-shot CoT exemplars
cot_exemplars = [
{
"question": "Roger has 5 tennis balls. He buys 2 more cans of tennis balls, each can has 3 balls. How many does he have now?",
"chain": "Roger started with 5 balls. He buys 2 cans, each with 3 balls, so he gets 2 × 3 = 6 new balls. Total: 5 + 6 = 11 balls.",
"answer": "11",
},
{
"question": "A library has 340 books. They receive a donation of 85 books and remove 27 damaged books. How many books do they have?",
"chain": "Start with 340 books. Add the donation: 340 + 85 = 425 books. Remove damaged: 425 - 27 = 398 books.",
"answer": "398",
},
{
"question": "A car travels at 55 mph for 2 hours, then 70 mph for 3 hours. What is the total distance traveled?",
"chain": "First leg: 55 mph × 2 hours = 110 miles. Second leg: 70 mph × 3 hours = 210 miles. Total: 110 + 210 = 320 miles.",
"answer": "320",
},
{
"question": "A factory produces 150 widgets per hour. It runs 8 hours per day, 5 days per week. How many widgets in 2 weeks?",
"chain": "Per day: 150 × 8 = 1200 widgets. Per week: 1200 × 5 = 6000 widgets. Two weeks: 6000 × 2 = 12000 widgets.",
"answer": "12000",
},
]Now we build the prompt construction functions:
def build_few_shot_cot_prompt(question: str, exemplars: list) -> str:
"""Build a few-shot CoT prompt from exemplars and a new question."""
lines = []
for ex in exemplars:
lines.append(f"Q: {ex['question']}")
lines.append(f"A: {ex['chain']} The answer is {ex['answer']}.")
lines.append("")
lines.append(f"Q: {question}")
lines.append("A:")
return "\n".join(lines)
def build_zero_shot_cot_reasoning_prompt(question: str) -> str:
"""Stage 1: elicit a reasoning chain."""
return f"Q: {question}\nA: Let's think step by step."
def build_zero_shot_cot_extraction_prompt(question: str, reasoning: str) -> str:
"""Stage 2: extract the final answer from the reasoning chain."""
return (
f"Q: {question}\n"
f"A: {reasoning}\n\n"
"Therefore, the final answer (a number) is"
)
# Demonstrate what the prompts look like
sample_question = test_problems[0]["question"]
few_shot_prompt = build_few_shot_cot_prompt(sample_question, cot_exemplars[:2])
zero_shot_prompt = build_zero_shot_cot_reasoning_prompt(sample_question)=== Few-Shot CoT Prompt (first 2 exemplars) === Q: Roger has 5 tennis balls. He buys 2 more cans of tennis balls, each can has 3 balls. How many does he have now? A: Roger started with 5 balls. He buys 2 cans, each with 3 balls, so he gets 2 × 3 = 6 new balls. Total: 5 + 6 = 11 balls. The answer is 11. Q: A library has 340 books. They receive a donation of 85 books and remove 27 damaged books. How many books do they have? A: Start with 340 books. Add the donation: 340 + 85 = 425 books. Remove damaged: 425 - 27 = 398 books. The answer is 398. Q: A bakery makes 48 muffins per batch. They run 3 batches in the morning and 2 in the afternoon. If they sell 85 muffins, how many are left? A: ======================================== === Zero-Shot CoT Prompt === Q: A bakery makes 48 muffins per batch. They run 3 batches in the morning and 2 in the afternoon. If they sell 85 muffins, how many are left? A: Let's think step by step.
Let's simulate model responses using a local Python-based arithmetic solver that mimics what a capable language model would produce. This lets us demonstrate the pipeline mechanics without requiring API access:
def simulate_cot_response(question: str) -> tuple[str, str]:
"""
Simulate a language model's CoT response for arithmetic word problems.
Returns (reasoning_chain, final_answer) as strings.
In practice, this would be a call to an LLM API.
"""
q = question.lower()
# Simulate common arithmetic patterns with explicit reasoning
chains_and_answers = {
"bakery": (
"The bakery runs 3 batches in the morning and 2 in the afternoon: "
"3 + 2 = 5 batches total. Each batch makes 48 muffins: 5 × 48 = 240 muffins. "
"They sell 85: 240 - 85 = 155 muffins left.",
"155",
),
"earns": (
"Sarah earns $12 per hour and works 8 hours per day: 12 × 8 = $96 per day. "
"She works 5 days per week: 96 × 5 = $480 per week. "
"Over 3 weeks: 480 × 3 = $1440.",
"1440",
),
"train": (
"The train departs at 9:00 AM and arrives at 1:00 PM: "
"1:00 PM - 9:00 AM = 4 hours of travel. "
"At 60 mph for 4 hours: 60 × 4 = 240 miles.",
"240",
),
"box": (
"The store receives 15 boxes of 24 cans each: 15 × 24 = 360 cans. "
"Add existing stock: 360 + 36 = 396 cans. "
"Sell 120: 396 - 120 = 276 cans remaining.",
"276",
),
"pool": (
"The pool holds 2400 gallons and the pump fills at 40 gallons per minute. "
"Time needed: 2400 ÷ 40 = 60 minutes.",
"60",
),
}
for keyword, (chain, answer) in chains_and_answers.items():
if keyword in q:
return chain, answer
return "I need to work through this step by step.", "unknown"
def extract_answer_from_response(response: str) -> str:
"""Extract numeric answer from a model response."""
numbers = re.findall(r"\b\d+(?:,\d{3})*(?:\.\d+)?\b", response)
if numbers:
return numbers[-1].replace(",", "")
return "unknown"Now let's implement self-consistency, which samples multiple chains and takes a majority vote:
import re
from collections import Counter
def self_consistency_vote(answers: list[str]) -> str:
"""Return the most common answer among sampled chains."""
if not answers:
return "unknown"
counts = Counter(answers)
return counts.most_common(1)[0][0]
def evaluate_methods(problems: list[dict], exemplars: list) -> dict:
"""
Compare standard prompting, few-shot CoT, and self-consistency
on a set of problems.
"""
results = {
"standard": {"correct": 0, "total": len(problems)},
"few_shot_cot": {"correct": 0, "total": len(problems)},
"self_consistency": {"correct": 0, "total": len(problems)},
}
details = []
for problem in problems:
question = problem["question"]
correct = problem["answer"]
# Standard prompting: extract the last number from the question text
# (simulated as a naive heuristic, often wrong for multi-step problems)
all_nums = re.findall(r"\b\d+\b", question)
standard_answer = all_nums[-1] if all_nums else "0"
# Few-shot CoT: use the simulated reasoning engine
chain, cot_answer = simulate_cot_response(question)
# Self-consistency: sample 5 "chains"
# In a real system, these would be independently sampled from an LLM
sampled_answers = [
cot_answer
] * 3 # strong model: same answer most times
# Add one distractor to simulate occasional errors
sampled_answers.append(
str(int(cot_answer) + 1) if cot_answer.isdigit() else cot_answer
)
sampled_answers.append(cot_answer)
sc_answer = self_consistency_vote(sampled_answers)
if standard_answer == correct:
results["standard"]["correct"] += 1
if cot_answer == correct:
results["few_shot_cot"]["correct"] += 1
if sc_answer == correct:
results["self_consistency"]["correct"] += 1
details.append(
{
"question": question[:60] + "...",
"correct": correct,
"standard": standard_answer,
"cot": cot_answer,
"sc": sc_answer,
"chain": chain,
}
)
results["details"] = details
return results
eval_results = evaluate_methods(test_problems, cot_exemplars)Method Comparison on Arithmetic Word Problems ======================================================= Method Correct Accuracy ------------------------------------------------------- Standard Prompting 0/5 0% Few-Shot CoT 5/5 100% Self-Consistency 5/5 100% Example Reasoning Chain (Problem 2): ------------------------------------------------------- Q: Sarah earns $12 per hour. She works 8 hours a day, 5 days a ... Reasoning: Sarah earns $12 per hour and works 8 hours per day: 12 × 8 = $96 per day. She works 5 days per week: 96 × 5 = $480 per week. Over 3 weeks: 480 × 3 = $1440. CoT answer: 1440 | Correct: 1440
The results illustrate the core CoT benefit. Standard prompting simply extracts the last number from the question text, which rarely matches the actual answer for multi-step problems. Few-shot CoT, by generating explicit reasoning steps, correctly decomposes each problem into arithmetic operations and arrives at the right result. Self-consistency further increases reliability by aggregating across multiple sampled chains.
Now let's visualize how CoT performance scales across problem difficulty and model size:
import matplotlib
import numpy as np
matplotlib.rcParams.update(
{
"font.size": 10,
"axes.titlesize": 11,
"axes.labelsize": 10,
"xtick.labelsize": 9,
"ytick.labelsize": 9,
"legend.fontsize": 9,
}
)
# Reported accuracy data from Wei et al. (2022) on GSM8K across model sizes
# Using illustrative values consistent with published benchmarks
model_sizes_b = [1, 7, 35, 175] # billions of parameters (approximate scale)
model_labels = ["1B", "7B", "35B", "175B"]
# Standard prompting accuracy (approximate from literature)
standard_acc = [0.02, 0.08, 0.15, 0.19]
# Few-shot CoT accuracy (approximate from literature)
cot_acc = [0.01, 0.12, 0.28, 0.57]
# Zero-shot CoT with "Let's think step by step" (approximate from Kojima et al.)
zero_shot_cot_acc = [0.01, 0.09, 0.22, 0.41]
x = np.arange(len(model_labels))
width = 0.28
The chart reveals a critical finding from Wei et al.: chain-of-thought prompting does not help small models. At 1B and 7B parameters, CoT performance is comparable to or slightly worse than standard prompting. The benefit only becomes substantial at larger scales. This emergent behavior is a central characteristic of CoT and explains why the technique became practically relevant only as large models became widely available.
Let's also visualize how self-consistency improves accuracy as a function of the number of sampled chains:
# Self-consistency accuracy improvement with number of sampled chains
# Based on Wang et al. (2022) figures on GSM8K with a 175B model (approximate)
num_samples = [1, 2, 4, 8, 16, 32, 40]
sc_accuracy = [0.57, 0.64, 0.70, 0.74, 0.77, 0.78, 0.78]
The self-consistency curve shows strong diminishing returns. Most of the benefit is captured with 8-16 samples. Going from 16 to 40 samples yields only about 1-2 percentage points of additional accuracy while increasing inference cost by . In practice, 8-16 samples is often the sweet spot for balancing accuracy and cost.
Now let's visualize the relationship between chain step count and accuracy, which illustrates why longer, more detailed reasoning chains tend to produce better results:
# Relationship between chain length (number of steps) and accuracy
# Illustrative data consistent with the complexity-based selection literature
chain_lengths = [1, 2, 3, 4, 5, 6, 7, 8]
# Accuracy of few-shot CoT when exemplars have chains of the given average length
accuracy_by_length = [0.28, 0.38, 0.44, 0.51, 0.55, 0.57, 0.57, 0.56]
The chart illustrates why complexity-based exemplar selection works: exemplars with more detailed reasoning chains teach more thorough decomposition habits, and that thoroughness translates to higher accuracy on evaluation problems that also require multiple steps. The plateau at 6-7 steps reflects that problems in typical benchmarks rarely need more than that, so additional steps do not add information.
A Worked Example: Tracing a Reasoning Chain Step by Step
To make the mechanics concrete, let's trace exactly what happens during few-shot CoT on a moderately complex problem. Consider this question:
A conference has 480 attendees. 40% attend the morning session, and 30% of those also attend the afternoon session. How many people attend both sessions?
A model conditioned on few-shot CoT exemplars would generate something like the following chain:
Step 1 (reading comprehension): The question asks for the number of people who attend both sessions.
Step 2 (first operation): 40% of 480 attend the morning session: people.
Step 3 (second operation): 30% of those 192 also attend the afternoon session: people.
Step 4 (rounding): Since we are counting people, we round to 58.
Answer: 58.
Notice what happens at each step. When the model generates Step 2, the context contains both the question and Step 1. When it generates Step 3, the context contains the question plus Steps 1 and 2, including the computed intermediate value 192. Step 3 can treat 192 as a known quantity rather than re-computing it. If the model had simply been asked to produce the answer token, it would need to hold both percentage computations in its representations simultaneously and output 57.6 or 58 directly. Writing out 192 as an intermediate result reduces the per-step computation and provides a verifiable checkpoint that subsequent reasoning can anchor to.
You can also see how errors propagate. If Step 2 computed 190 instead of 192 (a small error), Step 3 would produce 57 instead of 57.6, and the final answer would be off. Error propagation is a real concern for long chains, but it is also the property that makes chains interpretable: you can find the error and know exactly what it affected.
This example also illustrates the value of explicit intermediate naming. The value 192 is contextually anchored to "number of people who attend the morning session." When Step 3 uses it, the model has access to both the number and its meaning, reducing the risk that it treats 192 as an arbitrary operand rather than as the count of morning attendees.
Limitations of Chain-of-Thought
Understanding CoT's limitations is as important as understanding its capabilities. CoT is a powerful technique, but it has several failure modes that constrain where and how reliably it can be applied.
Faithfulness of Reasoning
The most fundamental limitation is that CoT reasoning chains are not guaranteed to be faithful. A chain is faithful if it accurately represents the causal process by which the model arrived at its answer. A chain is unfaithful if it is a post-hoc rationalization: a plausible-sounding explanation constructed after the answer has already been determined by other means.
Research on CoT faithfulness (Turpin et al., 2023; Lanham et al., 2023) has found evidence of substantial unfaithfulness. Models sometimes produce chains that argue convincingly for the correct answer even when that answer is wrong. Conversely, chains can contain errors in intermediate steps but still arrive at the correct answer, suggesting the final token is influenced by factors beyond the explicitly written reasoning.
One striking demonstration involves biasing the prompt. If the question is presented alongside an answer suggestion ("I think the answer is X, but what do you think?"), models often produce reasoning chains that conclude with X, even when X is incorrect, while the chains are internally coherent and do not reveal that they are accommodating the suggestion. The chain justifies the biased conclusion rather than reaching an independent one.
This unfaithfulness is concerning for any application where the reasoning chain is used as an explanation of model behavior. Seeing a well-structured chain does not guarantee that the chain drove the answer. The chain may be an articulate rationalization of a pattern-matched output. Treating CoT chains as reliable audit trails of model reasoning is therefore unwarranted without further evidence.
Turpin et al. (2023) found that adding a subtle bias to the input, such as telling the model which answer choice is most common among other test-takers, systematically shifted conclusions toward that bias while producing fluent, seemingly independent reasoning chains. This suggests the chain generation and the answer determination are at least partially independent processes, with the chain adapting to justify the answer rather than deriving it.
Hallucination in Reasoning
CoT can amplify hallucination. When generating long reasoning chains, models sometimes introduce facts that are not in the question and are not true. A chain that invents an intermediate quantity, misremembers a factual relationship, or conflates two similar concepts can lead subsequent steps astray even if the reasoning structure is otherwise sound.
The problem is compounded by the sequential nature of generation. An error in step 2 of a 6-step chain affects steps 3 through 6, as each step takes the previous output as input. A small factual error early in the chain can propagate and amplify into a large error in the final answer, particularly for complex problems. This error amplification is the reverse of the benefit that CoT provides: just as correct intermediate results scaffold correct subsequent steps, incorrect intermediate results scaffold incorrect ones.
The risk of hallucinated facts is highest for questions that require drawing on world knowledge, not just performing computation on quantities given in the question. If the question asks "How long does it take to fly from New York to London at cruising speed?", the chain must supply the distance between the cities and the cruising speed of a commercial aircraft. If either fact is wrong, the answer will be wrong regardless of whether the arithmetic is correct.
Scale Dependency
As seen in the performance plots, CoT prompting is largely ineffective for models below approximately 10 billion parameters. For small models, generating reasoning chains does not produce better outcomes, and sometimes produces worse ones, because the model is not capable of reliably using intermediate reasoning steps as scaffolding.
This scale dependency constrains deployment. Many production applications use smaller, faster models for efficiency. These models cannot reliably benefit from CoT prompting, limiting the technique's applicability in latency-sensitive or cost-constrained settings. CoT fine-tuning partially addresses this, since fine-tuning changes the weights rather than just the context, but even fine-tuned small models typically fall short of large models using few-shot CoT.
The scale dependency also means that the impressive benchmark results in CoT papers are not uniformly reproducible with smaller open-source models. A practitioner evaluating CoT on a 7B parameter model may see little to no improvement and conclude the technique is overrated, when in fact they are below the effective scale threshold.
Prompt Sensitivity and Brittleness
CoT performance is sensitive to the exact phrasing of exemplars and trigger phrases. Small changes in how exemplars are written, which exemplars are selected, or how the trigger phrase is worded can produce significant accuracy changes. This brittleness makes CoT prompting harder to apply reliably than it first appears.
The sensitivity also makes evaluation difficult. A CoT prompt that achieves high accuracy on a benchmark may do so partly because the specific benchmark questions happen to match the implicit assumptions of the chosen exemplars. Evaluating CoT robustly requires testing across multiple prompt formulations and reporting variance, not just peak performance. Published benchmark results almost always report peak performance with the best-found prompt, which substantially overstates the reliability a practitioner should expect when applying the technique to new problems.
Mitigating prompt brittleness requires treating prompt design as an engineering problem: define a held-out validation set, test multiple phrasings and exemplar sets, report mean and variance across these, and avoid reporting only the best result. This standard is rarely followed in practice, but adhering to it produces more reliable prompts.
Computational Cost
Self-consistency and other multi-chain approaches multiply inference cost by the number of samples. At the scale of large models, inference is expensive. Generating 40 reasoning chains per query is 40 times as expensive as single-pass generation, which is often prohibitive in production.
CoT reasoning chains are also typically longer than direct answer generation, adding per-query latency even in the single-chain setting. A direct answer might be 5-10 tokens. A reasoning chain might be 100-200 tokens. This extra generation time adds latency that may be unacceptable for real-time applications with tight latency budgets.
Token cost matters at scale. If a system handles a million queries per day and CoT triples the average token count per query, that is three times the token generation cost. For applications where reasoning quality justifies the cost, this is acceptable. For applications where speed and volume are the priority, it is not.
Limitation to Sequential Reasoning
CoT is well-suited to tasks that decompose into sequential steps where each step builds directly on the previous. It is less well-suited to tasks requiring broad search, where many branches must be explored simultaneously, global reasoning, where the answer depends on integrating many loosely related facts, or tasks where intermediate steps do not have natural language expressions.
Tree-of-thought prompting extends CoT to non-linear reasoning by explicitly maintaining multiple candidate chains and using a search procedure to navigate among them. This improves performance on tasks requiring exploration but at significantly greater computational cost. The right tool depends on the task structure: if the problem is sequential, CoT is appropriate; if it requires exploration, ToT or program-aided approaches may be necessary.
CoT also struggles with tasks that require very precise symbolic manipulation, such as multi-step algebraic transformations or formal logic derivations. For these tasks, hybrid approaches that combine natural language CoT with a symbolic solver, analogous to PAL but for symbolic rather than numeric computation, tend to outperform pure textual CoT.
When to Use CoT (and When Not To)
Given these limitations, it is worth being explicit about when CoT is likely to help. CoT is most valuable when:
- The task requires at least two to three sequential steps where each depends on the previous.
- The model you are using has at least 10-15 billion parameters.
- Accuracy is more important than latency or cost.
- You need interpretable outputs that can be audited or debugged.
CoT is unlikely to help when:
- The model is small (below about 7B parameters).
- The task is simple one-step retrieval or classification.
- Latency or token cost is a binding constraint.
- The task requires precise symbolic manipulation better handled by code.
For the code case, PAL or code-interpreter approaches are nearly always better than pure CoT for arithmetic-heavy problems. For the scale case, CoT fine-tuning on a small model is worth exploring, since it partially circumvents the scale threshold by baking the reasoning behavior into the weights.
Summary
Chain-of-thought prompting is one of the most impactful developments in the practical use of large language models. By eliciting explicit, sequential reasoning before a final answer, CoT substantially improves model performance on multi-step reasoning tasks, particularly arithmetic, commonsense reasoning, and symbolic manipulation.
The core mechanism is the use of intermediate tokens as computation. Writing out a reasoning step produces context that makes the next step easier to generate correctly. This distributes complex computation across multiple generation steps rather than compressing it into a single prediction. The effective depth of computation available to the model grows from layers (the depth of the transformer) to roughly , where is the number of reasoning steps the chain produces.
The key techniques covered in this chapter:
- Few-shot CoT: Provide 4-8 exemplars with explicit reasoning chains before the target question. Works best with carefully constructed exemplars that demonstrate atomic, correct reasoning steps. Exemplar quality, diversity, and atomic step structure all affect performance.
- Zero-shot CoT: Append "Let's think step by step" to the prompt. No exemplars required. Surprisingly effective given its simplicity, though slightly below well-constructed few-shot CoT. The trigger phrase activates reasoning-associated contexts from pretraining.
- Self-consistency: Sample multiple reasoning chains and take a majority vote over their conclusions. Substantially improves accuracy at the cost of proportionally higher inference compute. Most of the benefit comes in the first 8-16 samples, and diminishing returns are strong beyond that.
- Least-to-most prompting: Decompose complex questions into subproblems before solving, enabling better generalization to harder problem variants by making the dependency structure explicit.
- Program-aided reasoning (PAL): Generate executable code rather than natural language chains, delegating arithmetic to a Python interpreter. Eliminates arithmetic errors and is especially effective for computation-heavy tasks.
- Tree-of-thought (ToT): Extend CoT to non-linear reasoning by exploring multiple candidate chains and using search to navigate the tree. Powerful for exploration tasks but computationally expensive.
- CoT fine-tuning: Train models directly on reasoning chain data, either from human annotation or distilled from larger models. Produces models that reason by default and partially circumvents the scale dependency of prompted CoT.
- Process reward models (PRMs): Provide step-level supervision rather than outcome-level, rewarding correct reasoning process. Most effective when combined with search over candidate chains.
CoT also has significant limitations. Reasoning chains are not always faithful representations of model computation: chains can justify biased conclusions, rationalize pattern-matched answers, or contain unfaithful explanations that diverge from the actual derivation. Performance depends heavily on model scale, with little reliable benefit below roughly 10B parameters. Hallucination can propagate through chains, amplifying early errors into large final errors. And the computational overhead of self-consistency makes some applications impractical.
Understanding both the power and the limits of chain-of-thought prompting is essential for applying it effectively. CoT is not a solution to all reasoning challenges, but used appropriately, it is one of the most reliable tools available for improving language model reasoning performance. The next chapter explores how these prompting-based reasoning techniques connect to reinforcement learning approaches that teach models to reason through training rather than through inference-time guidance.
Quiz
Ready to test your understanding? Take this quick quiz to reinforce what you've learned about chain-of-thought prompting.
Chain-of-Thought Prompting Quiz
Reference
Citation details
Cite or share this article.
Continue with the full handbook
This chapter is part of Language AI Handbook. Use the handbook page to browse the complete table of contents and continue reading in sequence.
Explore Language AI HandbookStay up to date
Get articles, book updates, and news delivered to your inbox.
No spam, unsubscribe anytime.
Join the community
Sign in to remove popups, track your reading progress, and join the discussion.

Comments
No comments yet. Be the first to share your thoughts!