SMPC Demonstrations Enable Sparse RL for Loco-Manipulation
A new research paper on arXiv (2608.12063) introduces a method to train robots for combined locomotion and manipulation tasks using sparse rewards, bypassing the need for dense reward shaping. The approach leverages Sample-based Model Predictive Control (SMPC) entirely in simulation to generate massive offline datasets, solving the exploration problem. An off-policy RL agent is then trained with purely sparse task rewards, reducing learning time and eliminating manual tuning. The high-level agent is integrated with a low-level dynamic stability controller, yielding optimal behaviors aligned with task objectives, and the learned policies can surpass the original SMPC teacher. The framework's robustness is validated through sim-to-real transfer, though the abstract is cut off mid-sentence.
Key facts
- Paper ID: arXiv:2608.12063
- Announce Type: cross
- Method uses SMPC in simulation to generate offline datasets
- Trains off-policy RL with sparse rewards
- Integrates high-level agent with low-level stability controller
- Learned policies can surpass the optimal control teacher
- Validates sim-to-real robustness
- Aims to scale RL for complex loco-manipulation tasks
Entities
Institutions
- arXiv