Multi-Agent Debate
Multi-agent debate asks several model participants to propose answers, examine one another's arguments and revise their positions before a final decision. It is an inference-time technique for exposing alternative reasoning, with value assessed by correctness and evidence quality rather than the participants' eventual agreement.
What it is
A debate protocol specifies initial proposals, who can inspect which responses, the number of discussion rounds and how the final answer is selected. Participants may critique assumptions, identify missing steps or offer competing interpretations. Some implementations use the same model with different contexts, while others vary models or assigned roles. This differs from simple majority voting, which aggregates answers without an exchange, and from independent review against an external test. Debate changes the context available to each participant; it does not inherently add new factual evidence or update the models' parameters.
What the work involves
The practitioner chooses problems where disagreement can reveal a meaningful defect and defines a bounded exchange protocol. Initial responses should remain independent when diversity is important. Critiques need to address specific evidence or reasoning steps. A useful experiment compares debate with matched-budget independent sampling and a single-model baseline. Traces retain the original positions and revisions so evaluators can detect persuasion without correction. Final selection can use external verification when the task offers a checkable answer.
Illustrative example
Three agents analyze a scheduling puzzle independently, then inspect the constraints omitted by other participants. One identifies that two events must occupy different days, invalidating an early proposal. After one revision round, a deterministic constraint checker evaluates the candidate schedule. If all agents converge on an invalid arrangement, the checker rejects it. The example tests whether discussion repairs reasoning, rather than treating unanimous prose as sufficient proof.
Limits and common mistakes
Participants can share a misconception or be persuaded by confident but incorrect arguments. More rounds may homogenize answers and increase cost without improving accuracy. Role prompts do not necessarily create independent expertise. Quality requires evaluation of final correctness, informative disagreement and the cost of the exchange. Debate is poorly suited to resolving facts that require unavailable evidence; argument alone cannot replace a source lookup, calculation or experiment needed to establish the disputed claim.
Prerequisites
Debate is a multi-agent protocol.
- softSelf-Consistency
Both aggregate multiple reasoning attempts into one answer.
Sources and further reading
- Improving Factuality and Reasoning in Language Models through Multiagent Debate
Introduces multi-round exchanges among language-model instances and evaluates their effect on reasoning and factual tasks.
Last updated: 2026-10-10