Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
A novel training approach for human-in-the-loop (HIL) online reinforcement learning in actual robots has been detailed in a paper on arXiv (2608.15088). This technique integrates two key elements: an MC Q-chunk critic and max-Q selective imitation. The critic evaluates chunk-level action values using Monte Carlo returns from the replay buffer, enabling direct crediting of intervention trajectories through sample-average policy assessment. The actor updates by mimicking the higher-Q action at each state, choosing between the current policy action and a buffer sample based on a strict winner-take-all principle. This method facilitates a swift integration of human interventions while fostering ongoing improvement beyond the initial human input. The research team is credited as the authors, and the paper can be found on arXiv.
Key facts
- The paper is available on arXiv with ID 2608.15088.
- The method uses an MC Q-chunk critic for sample-average policy evaluation.
- Max-Q selective imitation uses a hard winner-take-all rule.
- The rule switches between learning from interventions and on-policy self-improvement.
- The method is designed for real robots in human-in-the-loop online reinforcement learning.
- It aims to absorb human interventions quickly while improving beyond the human prior.
- The approach reduces the policy-target-sample gap.
- The paper is a cross-type announcement.
Entities
Institutions
- arXiv