ARTFEED — Contemporary Art Intelligence

SOOPER: Safe RL via Conservative Policy Priors

ai-technology · 2026-08-15

A novel reinforcement learning (RL) algorithm, SOOPER (Safe Optimistic Exploration using Policy Priors), has been developed to tackle the issue of safe exploration for RL agents. This approach utilizes conservative, albeit suboptimal, policies that can be obtained from offline datasets or simulations as priors. By employing probabilistic dynamics models, SOOPER facilitates optimistic exploration while ensuring a conservative policy serves as a fallback when needed. The researchers have demonstrated theoretically that SOOPER maintains safety during the learning phase and ensures convergence to an optimal policy by limiting cumulative regret. Comprehensive testing on significant safe RL benchmarks and real-world hardware confirms SOOPER's scalability and superiority over leading methods, affirming its theoretical claims. This paper falls under Computer Science > Machine Learning and is available on arXiv with the identifier 2601.19612, highlighting its importance for RL systems that must function safely in real-world scenarios to prevent adverse effects from uncontrolled exploration.

Key facts

  • SOOPER is a new RL algorithm for safe exploration.
  • It uses suboptimal yet conservative policies as priors.
  • Probabilistic dynamics models enable optimistic exploration.
  • Pessimistic fallback to the conservative policy ensures safety.
  • Theoretical guarantees include safety and convergence to optimal policy.
  • Cumulative regret is bounded.
  • Experiments on safe RL benchmarks and real-world hardware show scalability and outperformance.
  • Paper submitted to arXiv with ID 2601.19612.

Entities

Institutions

  • arXiv

Sources