LEMUR: Multi-Objective RL from Preference Feedback
A new arXiv preprint (2607.29559) introduces LEMUR (Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback), a method that combines multi-objective reinforcement learning (MORL) with preference-based RL (PbRL). Traditional RL relies on a single scalar reward, but real-world tasks often involve competing objectives like performance vs. efficiency, where specifying reward functions is difficult. MORL addresses trade-offs with vector rewards but still assumes well-specified reward functions. PbRL learns rewards from human feedback, avoiding predefined rewards, but has been limited to single-objective settings. LEMUR bridges this gap by enabling RL agents to learn from preference feedback in multi-objective scenarios, potentially improving alignment with human values in complex decision-making tasks. The paper is announced as a new submission and is available on arXiv.
Key facts
- arXiv:2607.29559
- Announce Type: new
- LEMUR stands for Learning to Align with Multi-Objective Reinforcement Learning with Preference feedback
- Combines Multi-Objective RL (MORL) and Preference-based RL (PbRL)
- Addresses challenges of specifying reward functions in multi-objective tasks
- Real-world tasks often involve competing objectives such as performance vs. efficiency
- Existing MORL approaches assume well-specified reward functions
- PbRL learns rewards from human feedback but has been studied in single-objective settings
Entities
Institutions
- arXiv