Hybrid Attention Estimation Pipeline for Adaptive Human-Robot Interaction
A study published on arXiv (2608.00284) explores hybrid visual attention estimation for human-robot interaction, utilizing an expressive robotic head from the InMoov ecosystem. This innovative approach merges a rapid geometric perception layer with a separate semantic perception layer that employs a vision-language model. The geometric layer delivers high-frequency data on face and head poses for temporal adjustments, while the semantic layer processes raw egocentric camera frames to generate contextual attention labels, indicating focus on the robot, phone usage, or other distractions. A finite state machine integrates these signals to manage adaptive interaction behaviors, such as activation, waiting, resuming interaction, and returning to rest. The system underwent evaluation with 10 participants in 40 trials, demonstrating consistent interaction initiation across all scenarios.
Key facts
- Paper arXiv:2608.00284
- Hybrid visual attention estimation pipeline
- Expressive robotic head based on InMoov ecosystem
- Combines geometric and semantic perception layers
- Vision-language model for semantic layer
- Finite state machine for adaptive interaction
- Evaluated with 10 participants and 40 trials
- Results show reliable interaction start
Entities
Institutions
- arXiv
- InMoov