New arXiv Paper on Adaptive Policy Portfolios for Robust Markov Decision Processes
A recent paper on arXiv (2608.17929v1) presents adaptive policy portfolios tailored for robust Markov decision processes (RMDPs). Conventional RMDPs typically choose a single policy that performs adequately across various plausible transition functions, which may lead to excessive caution when dynamics remain unknown yet partially identifiable post-deployment. The authors suggest generating a finite collection of memoryless randomized policies offline, complemented by a simple online selector. They introduce robust regret as a metric for evaluating portfolio effectiveness: it measures the loss of the top portfolio member compared to the optimal policy for a known environment. This research expands on regret objectives explored by Ghavamzadeh et al. (2016), emphasizing safe policy enhancement. Additionally, the paper offers a complexity-theoretic perspective, demonstrating that certifying a portfolio is ∀R-complete, even for deterministic portfolios in acyclic (s,a)-rectangular RMDPs.
Key facts
- The paper is arXiv:2608.17929v1, announced as a new submission.
- It studies robust Markov decision processes with adaptive policy portfolios.
- Portfolios are finite sets of memoryless randomized policies synthesized offline.
- A lightweight online selector is used to choose among portfolio members.
- Robust regret is introduced as a measure of portfolio quality.
- Related regret objectives by Ghavamzadeh et al. (2016) are referenced.
- Certifying a portfolio is proven to be ∀R-complete for deterministic portfolios in acyclic (s,a)-rectangular RMDPs.
- The approach aims to reduce conservatism when unknown dynamics become partially identifiable.
Entities
Artists
- Mohammad Ghavamzadeh
Institutions
- arXiv