Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry
A recent study in computer science tackles the multi-armed bandit issue in scenarios with heavy-tailed reward distributions and decentralized interactions characterized by information asymmetry. The researchers introduce resilient decentralized algorithms applicable to three scenarios: common rewards with unobserved actions, independent rewards with observed actions, and independent rewards with unobserved actions. They establish regret guarantees that closely align with centralized heavy-tailed rates. Experiments conducted in a Pareto-distributed reward setting support the theoretical conclusions and highlight the trade-offs involved in synchronization, coordination, and exploration. This paper is accessible on arXiv (2608.10529) and is categorized under machine learning.
Key facts
- The paper studies multi-agent multi-armed bandits with heavy-tailed rewards.
- Three information-asymmetry regimes are considered: unobserved actions with common rewards, observed actions with independent rewards, and unobserved actions with independent rewards.
- Robust decentralized algorithms are developed for each setting.
- Regret guarantees nearly match centralized heavy-tailed rates.
- Experiments on a Pareto-distributed reward environment validate theoretical findings.
- The paper is available on arXiv with ID 2608.10529.
- The paper is categorized under Computer Science > Machine Learning.
- The paper discusses trade-offs between synchronization, coordination, and exploration.
Entities
Institutions
- arXiv