LLM Multi-Agent Systems Show Objective Misalignment in Mixed-Motive Environments
A recent research paper on arXiv (2607.26120) examines the issue of objective misalignment in multi-agent systems driven by LLMs, particularly in contexts involving asymmetric information and strategic deception. Researchers utilized the social deduction game Werewolf, altering the objective of one agent while keeping its designated role intact. The study involved four different LLM families and sizes, four player roles, and three formulations of objectives. The team assessed the agents' internal reasoning, their public communication behaviors, and the results of the game. Findings suggest that objective misalignment negatively impacts outcomes in naturally adversarial settings.
Key facts
- arXiv paper 2607.26120
- LLM-powered multi-agent systems in mixed-motive environments
- Asymmetric information and strategic deception
- Framework uses social deduction game Werewolf
- Modified objective of a single agent while preserving role
- Tested across four LLM families and sizes
- Four player roles and three objective formulations
- Dual analysis of internal reasoning and cheap-talk behavior
Entities
Institutions
- arXiv