Atlas · skill

Self-Consistency

Self-consistency samples several reasoning paths and aggregates their final answers instead of trusting one generation. The method seeks a result supported by multiple sampled paths, commonly through majority voting, while recognizing that agreement among outputs from the same model is different from independent verification of the answer.

conceptReasoning Techniques

What it is

The original technique combines chain-of-thought prompting with stochastic sampling. Each sample develops a path and produces an answer; an aggregation rule then chooses the answer appearing most often after normalization. The mechanism relies on different valid paths converging on a result while some errors diverge. It is a decoding and aggregation procedure rather than a training update. Variants may use weighted votes or other selectors, but those choices need explicit definition. For open-ended outputs, deciding when two answers are equivalent is harder than comparing a number or a class label.

What the work involves

The practitioner defines an answer extraction rule, sampling settings, number of trials and aggregation behavior before evaluating results. Invalid or missing answers require a policy, and ties may trigger a fallback or more sampling. Test cases compare single-generation quality with aggregated quality under the same budget constraints. The artifact should retain sampled outputs and the normalized votes so failures are inspectable. An external correctness check is especially useful where the model can repeatedly reproduce the same mistaken assumption across apparently different explanations.

Illustrative example

A model solves a collection of arithmetic word problems. Several samples produce different explanations but the same numerical result, while others misread a quantity. The application extracts a normalized number from each sample and takes the most frequent valid answer. For a problem with ambiguous units, the votes may split or all repeat one interpretation. Those cases are routed for further checking rather than describing the vote margin as a calibrated probability that the answer is true.

Limits and common mistakes

Repeated model samples share training, prompts and often the same blind spots. A majority can therefore be wrong, especially when the question invites a common misconception. Sampling multiplies inference cost and latency, and aggregation can hide disagreement if answer normalization is too coarse. The technique should be evaluated on the target task against other ways to spend the same budget, including a stronger model, executable verification or better evidence retrieval.

Prerequisites

Sources and further reading

Last updated: 2026-10-10