RCI Framework Enhances Sparse Safe Offline RL with Redistribution-based Cost Inference
The Redistribution-based Cost Inference (RCI) framework tackles a significant issue in safe offline reinforcement learning (RL): the dependence on detailed per-step cost annotations. Typically, supervisors only give trajectory-level stop-feedback, a binary indication of the first unsafe transition without specific per-step details. The framework, outlined in arXiv:2608.12306, transforms this limited feedback into comprehensive per-step costs through return decomposition, facilitating the training of constrained offline policies on enhanced datasets. The study reveals that return-equivalent redistribution maintains the feasible policy set and optimal Lagrangian in a Constrained Markov Decision Process (CMDP), proving the transformation to be theoretically lossless while enhancing cost critic learning. Experiments in highway driving and robotic manipulation indicate significantly reduced violation rates compared to sparse and classifier-based methods, demonstrating robustness against varied dataset compositions. This research is crucial for practical applications where detailed cost annotations are unfeasible, providing a viable approach for safe RL in intricate environments.
Key facts
- RCI converts sparse stop-feedback into dense per-step costs via return decomposition.
- The framework preserves the feasible policy set and optimal Lagrangian in a CMDP.
- Experiments on highway driving and robotic manipulation show lower violation rates.
- RCI outperforms sparse and classifier-based baselines.
- The method is robust to heterogeneous dataset composition.
- The paper is available on arXiv with ID 2608.12306.
- The approach addresses temporal credit assignment in safe offline RL.
- The transformation is theoretically lossless.
Entities
Institutions
- arXiv