Preference Learning Recast as Human-Autonomy Team Problem
A new arXiv paper (2608.11229) proposes a paradigm shift in preference-based reward learning for human-robot interaction. The authors argue that the traditional approach, which treats the human teacher as a passive oracle answering learner-generated queries, is suboptimal. Instead, they recast preference learning as a human-autonomy team problem, where the teacher actively constructs training examples based on a model of the learner's current knowledge. This approach, termed second-order theory-of-mind, enables the teacher to design an informative curriculum that is more efficient than learner-driven acquisition, especially as the feature dimension of the reward grows. The paper extends this framework to include a learner that maintains a second-order model of the teacher, creating a coupled system. The work is relevant to AI, robotics, and human-robot collaboration, and was announced as a new arXiv preprint.
Key facts
- Paper arXiv:2608.11229v1, announced as new
- Proposes recasting preference learning as human-autonomy team problem
- Teacher uses model of learner to design curriculum
- Learner maintains second-order model of teacher
- Advantage widens with feature dimension growth
- Critiques passive oracle approach
- Published on arXiv
- Relevant to AI and robotics
Entities
Institutions
- arXiv