UC-PSRO: AI for Game-Theoretic Course-of-Action Generation in Adversarial Swarms
A new study on arXiv (2608.15372) introduces UC-PSRO, a method aimed at developing game-theoretically sound Courses of Action (COAs) for a Blue UAS swarm dealing with an adaptive Red adversary in settings where communication isn't reliable. While it draws inspiration from a U.S. Air Force SBIR request, the work stands on its own. UC-PSRO combines three elements: using PSRO self-play to train strategies for both sides, FiLM conditioning for adapting the Blue strategy based on a Commander's-Intent weight vector from a Dirichlet distribution, and a curriculum that encourages decentralized communication within the swarm. The research explores this approach in a specific scenario, addressing challenges in coordinating drone swarms under poor communication conditions, which is crucial for modern military missions.
Key facts
- Paper arXiv:2608.15372, announced as new
- Proposes UC-PSRO (Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum)
- Generates game-theoretically optimized Courses of Action (COAs) for Blue UAS swarm against adaptive Red adversary
- Motivated by public U.S. Air Force SBIR solicitation
- Combines PSRO self-play, FiLM conditioning on Commander's-Intent weight vector, and curriculum annealing of communication-graph edge dropout
- Blue policy is re-steerable at execution time without retraining
- Swarm learns decentralized, peer-to-peer fallback instead of depending on full connectivity
- Evaluation details are incomplete in the source
Entities
Institutions
- U.S. Air Force