Study Reveals Single Frontier Governing Conformity in LLMs
A new preprint on arXiv (2608.11247) investigates how large language models (LLMs) conform to peer opinions in collaborative multi-agent settings. The study, conducted across 23 open-weight models, 19 conditions, and three datasets, generated over a million graded responses. It found that a unanimous wrong majority reverses 22.8% of correct MMLU answers, 54.8% on GPQA, and 71.0% on SimpleQA, with 84-89% of reversed answers matching the peers' answers. The authors propose that existing mitigations focus solely on Resistance (keeping correct answers under pressure), but collaborating agents also need Receptivity (the rate at which a model accepts correct peer answers). They introduce a Resistance-Receptivity frontier, suggesting that improvements in one dimension often come at the cost of the other. The research highlights the vulnerability of LLMs to social conformity and the need for balanced mitigation strategies.
Key facts
- Study on arXiv:2608.11247
- 23 open-weight models tested
- 19 conditions and three datasets used
- Over one million graded responses
- Unanimous wrong majority reverses 22.8% of correct MMLU answers
- 54.8% reversal on GPQA
- 71.0% reversal on SimpleQA
- 84-89% of reversed answers match peers' answers
- Introduces Resistance-Receptivity frontier
Entities
Institutions
- arXiv