StaQ: Finite Memory Policy Mirror Descent for Discrete Action Spaces
A new arXiv paper proposes StaQ, a Policy Mirror Descent (PMD) algorithm for discrete action spaces that retains only the last M Q-functions in memory, addressing the intractability of exact PMD. The authors demonstrate theoretically that for a finite and sufficiently large M, the algorithm achieves convergence guarantees, providing a practical approximation that mitigates policy evaluation errors. The work, released as arXiv:2506.13862v2, contributes to reinforcement learning by offering a memory-efficient alternative to natural policy gradient and actor-critic methods, which can introduce errors. The paper is available on arXiv and is relevant to researchers in AI and machine learning.
Key facts
- Paper title: StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent
- arXiv ID: 2506.13862v2
- Announce type: replace-cross
- Proposes PMD-like algorithms for discrete action spaces keeping only last M Q-functions in memory
- Shows theoretically that for finite and large enough M, an RL algorithm achieves convergence
- Addresses intractability of exact PMD which requires sum of all past Q-functions
- Compares to natural policy gradient and actor-critic approaches which may introduce errors
- Published on arXiv (https://arxiv.org/abs/2506.13862)
Entities
Institutions
- arXiv