dorsal/arxiv
View SchemaDeliberationBench: When Do More Voices Hurt? A Controlled Study of Multi-LLM Deliberation Protocols
| Authors | Vaarunay Kaushal, Taranveer Singh |
|---|---|
| Categories | |
| ArXiv ID | 2601.08835vv1 |
| URL | https://arxiv.org/abs/2601.08835 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Multi-agent systems where Large Language Models (LLMs) deliberate to form consensus have gained significant attention, yet their practical value over simpler methods remains under-scrutinized. We introduce DELIBERATIONBENCH, a controlled benchmark evaluating three deliberation protocols against a strong baseline of selecting the best response from a pool of model outputs. Across 270 questions and three independent seeds (810 total evaluations), we find a striking negative result: the best-single baseline achieves an 82.5% +- 3.3% win rate, dramatically outperforming the best deliberation protocol(13.8% +- 2.6%). This 6.0x performance gap is statistically significant (p < 0.01) and comes at 1.5-2.5x higher computational cost. Our findings challenge assumptions that complexity enhances quality in multi-LLM systems.
{
"annotation_id": "9cbfe48b-1d2c-4865-a680-8d6aca0f0968",
"date_created": "2026-02-17T05:53:20.004000Z",
"date_modified": "2026-02-17T05:53:20.004000Z",
"file_hash": "42a27557cb994270f476ffffc1c2045404d6d84f8ea8cbfbe6e8583df957dc95",
"private": false,
"record": {
"abstract": "Multi-agent systems where Large Language Models (LLMs) deliberate to form consensus have gained significant attention, yet their practical value over simpler methods remains under-scrutinized. We introduce DELIBERATIONBENCH, a controlled benchmark evaluating three deliberation protocols against a strong baseline of selecting the best response from a pool of model outputs. Across 270 questions and three independent seeds (810 total evaluations), we find a striking negative result: the best-single baseline achieves an 82.5% +- 3.3% win rate, dramatically outperforming the best deliberation protocol(13.8% +- 2.6%). This 6.0x performance gap is statistically significant (p \u003c 0.01) and comes at 1.5-2.5x higher computational cost. Our findings challenge assumptions that complexity enhances quality in multi-LLM systems.",
"arxiv_id": "2601.08835",
"authors": [
"Vaarunay Kaushal",
"Taranveer Singh"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "DeliberationBench: When Do More Voices Hurt? A Controlled Study of Multi-LLM Deliberation Protocols",
"url": "https://arxiv.org/abs/2601.08835",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "62b8f39d-689e-4d37-a523-c2c0c05b4772",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}