ARTFEED — Contemporary Art Intelligence

Per-Agent Policy Composition in Multi-Agent RL Can Fail: New Proof

ai-technology · 2026-08-13

A new paper on arXiv (2608.11658) challenges the safety of per-agent policy composition in cooperative multi-agent reinforcement learning. The authors prove that when each agent independently recombines its own library of learned policies using successor features and generalized policy improvement, the resulting joint behavior can be strictly worse than every policy in the library. This is because recombining teammates changes the environment each agent faces, invalidating the values used for transfer. The paper highlights a critical gap in multi-agent transfer learning, which is common in applications like fleet management and traffic signal control where objectives change dynamically after deployment. The authors argue that while single-agent successor features with generalized policy improvement guarantee performance no worse than any policy in the library, the independent composition approach in multi-agent settings lacks such guarantees. The paper is a cross-type announcement and is available at https://arxiv.org/abs/2608.11658.

Key facts

  • Paper arXiv:2608.11658v1
  • Proves independent per-agent policy composition can be strictly worse than all policies in library
  • Focuses on cooperative multi-agent reinforcement learning
  • Applications include fleet management and traffic signal control
  • Single-agent successor features with generalized policy improvement guarantee no worse than library
  • Multi-agent transfer lacks similar guarantee
  • Recombining teammates changes environment each agent faces
  • Available at https://arxiv.org/abs/2608.11658

Entities

Institutions

  • arXiv

Sources