Wrong Answers Can Improve Multi-Agent Reasoning, Study Finds
A recent investigation published on arXiv (2608.14375) questions the prevalent belief in multi-agent reasoning systems that only accurate messages should affect the final outcomes. The authors present Diverse Hypothesis Deliberation (DHD), a systematic evaluation method that stores five independently created messages and replays them to a downstream solver (the integrator) with options to show or hide each message. This comparison assesses a message's 'trajectory value'—determining whether its availability aids or hinders reasoning. The research, covering five benchmarks in mathematics and science using two publicly available model families (gpt-oss-120b and gemma-4-31B-it), reveals that 'wrong-helpful' messages—incorrect yet beneficial answers—are present in all tested combinations. The findings suggest that such messages can enhance the integrator's reasoning, indicating that relying solely on correctness may overlook valuable insights. The study is titled 'Wrong but Useful: Trajectory Value Beyond Answer Correctness in Multi-Agent Messages.'
Key facts
- Study introduces Diverse Hypothesis Deliberation (DHD) protocol.
- DHD caches five independently generated messages and replays them with the integrator.
- Measures trajectory value: whether making a message available helps or harms reasoning.
- Tested on five mathematics and science benchmarks.
- Used two model families: gpt-oss-120b and gemma-4-31B-it.
- Wrong-helpful messages found in every benchmark-model combination.
- Challenges assumption that correct messages are the only valuable ones.
- Published on arXiv with ID 2608.14375.
Entities
Institutions
- arXiv