ARTFEED — Contemporary Art Intelligence

Preference Learning Recast as Human-Autonomy Team Problem

ai-technology · 2026-08-13

A new arXiv paper (2608.11229) proposes a paradigm shift in preference-based reward learning for human-robot interaction. The authors argue that the traditional approach, which treats the human teacher as a passive oracle answering learner-generated queries, is suboptimal. Instead, they recast preference learning as a human-autonomy team problem, where the teacher actively constructs training examples based on a model of the learner's current knowledge. This approach, termed second-order theory-of-mind, enables the teacher to design an informative curriculum that is more efficient than learner-driven acquisition, especially as the feature dimension of the reward grows. The paper extends this framework to include a learner that maintains a second-order model of the teacher, creating a coupled system. The work is relevant to AI, robotics, and human-robot collaboration, and was announced as a new arXiv preprint.

Key facts

  • Paper arXiv:2608.11229v1, announced as new
  • Proposes recasting preference learning as human-autonomy team problem
  • Teacher uses model of learner to design curriculum
  • Learner maintains second-order model of teacher
  • Advantage widens with feature dimension growth
  • Critiques passive oracle approach
  • Published on arXiv
  • Relevant to AI and robotics

Entities

Institutions

  • arXiv

Sources