Conquer-and-Merge Discussion (CMD) · ACL 2024
Qineng Wang1*, Zihao Wang2*, Ying Su2, Hanghang Tong3, Yangqiu Song2
1Zhejiang University 2HKUST 3UIUC
*Equal contribution
- [2026-09-28] Updated implementation with async OpenRouter inference and MindCube-Tiny / AIME results.
- [2024-08] Paper published at ACL 2024.
This is the official repository for Rethinking the Bounds of LLM Reasoning: Are Multi-Agent Discussions the Key? It implements CMD for math and visual reasoning, supporting single agents, same-model discussions, and discussions across different models.
CMD divides agents into groups. Agents exchange answers and explanations within their group and see answer counts from other groups. A final vote selects the answer; a secretary resolves ties.
flowchart LR
Q["Question + optional images"] --> I["Independent answers<br/>6 agents"]
I --> D["Group discussion<br/>2 groups of 3 · 2 updates"]
D --> V{"Final vote"}
V -->|Unique plurality| A["Answer"]
V -->|Tie| S["Secretary"]
S --> A
Requires Python 3.10+; no local GPU is needed.
git clone https://github.com/HKUST-KnowComp/LLM-discussion.git
cd LLM-discussion
python -m venv .venv
source .venv/bin/activate
python -m pip install -e '.[benchmarks]'Run a complete discussion offline with the included GSM8K data:
llm-discussion run --dry-run --scenario tie \
--dataset data/gsm8k-test.jsonl --limit 2 --output runs/quickstart
llm-discussion evaluate runs/quickstartDry runs need no API key and use synthetic responses, so they do not report benchmark accuracy.
For live evaluation, set OPENROUTER_API_KEY in your environment or a local .env file using .env.example. .env is ignored by Git.
llm-discussion prepare mindcube
llm-discussion suite --config examples/vision-suite.json \
--dataset data/benchmarks/mindcube/tasks.jsonl --limit 3 --sample-seed 42 \
--task-concurrency 12 --request-concurrency 24 --output runs/mindcube-live
llm-discussion evaluate runs/mindcube-liveThe suite runs three single-agent baselines, three homogeneous CMD conditions, and one mixed CMD condition through OpenRouter. Configure models in vision-suite.json or math-suite.json:
| Model | API reasoning |
|---|---|
openai/gpt-6-luna |
Disabled |
google/gemma-4-31b-it |
Disabled |
meta/muse-spark-1.3-contributor |
Minimal |
See the running guide for full evaluations and custom settings, or the data format to use your own questions.
September 2026 evaluation on MindCube-Tiny (1,050 questions), AIME 2025 (30), and AIME 2026 (30) using the models above. Each arrow shows single agent → homogeneous CMD accuracy (%).
| Model / condition | MindCube-Tiny | AIME 2025 | AIME 2026 |
|---|---|---|---|
| GPT-6 Luna | 46.76 → 48.67 | 16.67 → 33.33 | 36.67 → 40.00 |
| Gemma 4 | 57.05 → 60.29 | 56.67 → 63.33 | 53.33 → 66.67 |
| Muse Spark | 71.33 → 75.81 | 86.67 → 90.00 | 90.00 → 93.33 |
| Mixed CMD | 70.38 | 70.00 | 76.67 |
CMD uses six agents in two groups for three rounds. Mixed groups contain one agent per model, with Luna as secretary. Scores cover the full question sets; unsuccessful runs count as incorrect. These are results with current models and datasets.
python -m pip install -e '.[dev,benchmarks]'
python -m pytest -q
llm-discussion --helpImplementation: src/llm_discussion/. Configurations: examples/. Tests use synthetic agents and mocked APIs.
If you use CMD, please cite:
@inproceedings{wang-etal-2024-rethinking-bounds,
title = "Rethinking the Bounds of {LLM} Reasoning: Are Multi-Agent Discussions the Key?",
author = "Wang, Qineng and Wang, Zihao and Su, Ying and Tong, Hanghang and Song, Yangqiu",
booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
year = "2024",
publisher = "Association for Computational Linguistics",
pages = "6106--6131",
doi = "10.18653/v1/2024.acl-long.331",
url = "https://aclanthology.org/2024.acl-long.331/"
}