Chain-of-Thought Prompting
Chain-of-thought prompting asks a language model to produce intermediate reasoning steps before or alongside an answer. It can help some multi-step tasks by providing a pattern for decomposing the problem, but the visible explanation is generated text and should not be treated as a faithful record of the model's internal computation.
What it is
The original few-shot approach demonstrates problems with intermediate steps and final answers, encouraging the model to continue in the same style. A related zero-shot approach asks for stepwise reasoning without worked examples. Both modify the input and output behavior of a trained model rather than supplying a symbolic proof engine. The resulting steps can expose an arithmetic mistake or omitted condition, but they can also rationalize a wrong answer. Models specifically trained for reasoning may use provider-supported reasoning settings instead, so explicit requests for a detailed visible chain are not universally appropriate.
What the work involves
The practitioner tests whether intermediate steps improve the actual task, comparing direct answers and supported reasoning configurations under a recorded budget. Output checks focus on the final answer and externally verifiable steps rather than explanation length. For sensitive applications, a concise rationale tied to evidence may be more appropriate than a full generated chain. The artifact includes the prompting method, demonstrations if used, task scores and error analysis. Mathematical or code-based steps can be checked with tools when their correctness matters.
Illustrative example
An assistant solves a delivery-planning question requiring several durations to be combined. A worked example shows how to identify each duration and convert units before calculating the total. In testing, the approach reduces some unit mistakes but still fails when a waiting period overlaps travel. Reviewers inspect the final schedule against the stated conditions. A fluent sequence of steps that double-counts the overlap is marked wrong even when it looks more explanatory than a direct answer.
Limits and common mistakes
Visible reasoning can omit influential information, invent a justification or reproduce a shared misconception. Asking the model to reason does not guarantee improved accuracy, and additional tokens increase latency. Reported gains depend on the model, demonstrations and task distribution. A useful distinction is between an explanatory answer, a correct answer and a verified derivation: chain-of-thought prompting may assist the first two, while independent checks are needed for the third.
Prerequisites
Related skills
- → is subcategory of: Prompt Engineering
Sources and further reading
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Introduces prompting with intermediate reasoning demonstrations and evaluates selected reasoning tasks.
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting
Provides evidence that generated chains can misrepresent influences on a model's answer.
Last updated: 2026-10-10