Yokai Learning Environment: New Benchmark for Zero-Shot Coordination in Multi-Agent RL
The Yokai Learning Environment (YLE) has been unveiled by researchers as an open-source benchmark for multi-agent reinforcement learning, aimed at overcoming the shortcomings of the Hanabi Learning Environment (HLE). Historically, HLE has served as the primary benchmark for zero-shot coordination (ZSC), assessing algorithms by matching independently trained agents. However, recent developments have led to nearly flawless inter-seed cross-play performance in HLE, limiting its effectiveness in tracking algorithmic advancements. YLE incorporates new features not found in HLE, such as the ability to track and update beliefs regarding moving cards, reasoning with ambiguous hints, and determining when to end the game based on inferred shared knowledge. These capabilities challenge agents to establish common ground, a key aspect of cooperative AI. Leading ZSC methods, including High-Entropy IPPO, are evaluated within this benchmark. The paper can be accessed on arXiv with the identifier 2508.12480, categorized as 'replace'.
Key facts
- The Yokai Learning Environment (YLE) is introduced as a new benchmark for zero-shot coordination.
- YLE is open-source and designed for multi-agent reinforcement learning.
- The Hanabi Learning Environment (HLE) has been the dominant benchmark for ZSC.
- Recent work has achieved near-perfect inter-seed cross-play performance in HLE.
- YLE requires tracking and updating beliefs over moving cards.
- YLE includes reasoning under ambiguous hints.
- YLE requires deciding when to terminate the game based on inferred shared knowledge.
- Leading ZSC methods, including High-Entropy IPPO, are evaluated in YLE.
Entities
Institutions
- arXiv