Swap-Regret Loss in Single-Layer Self-Attention Models
A new study revisits the regret loss framework from Park et al. (2025), applying it to probability-simplex policies in single-layer self-attention models. The authors prove that training with regret loss yields a stationary point whose forward pass matches smoothed fictitious play, ensuring no-regret behavior. They also introduce a swap-regret loss function, extending the framework beyond external regret to optimize for swap-deviation robustness. This swap-regret loss admits a stationary point implementing the classical Blum-Mansour swap-regret update. The work bridges decision theory and machine learning, offering theoretical guarantees for model training.
Key facts
- arXiv:2607.23333v1
- Regret loss framework from Park et al. (2025)
- Single-layer self-attention model
- Probability-simplex policies
- Stationary point matches smoothed fictitious play
- New swap-regret loss function introduced
- Swap-regret loss enables swap-deviation robustness
- Stationary point implements Blum-Mansour update
Entities
Institutions
- arXiv