Pegasus Framework Bridges Embodiment Gap for Robot Learning
A novel framework named Pegasus, outlined in a preprint on arXiv (2607.26903), tackles the data bottleneck in embodied AI. Although there are billions of videos showcasing human manipulation, robots struggle to learn from them due to the disparity between human and robotic forms. Pegasus converts human demonstrations into data suitable for robots through structured knowledge transfer, utilizing a graph-based intermediate representation. It transforms a Task Graph from human videos into a Robot Planning Graph via Affordance and Constraint Graphs, facilitating robot-conditioned video generation. Additionally, a hierarchical affordance latent space captures object states, tasks, and affordances, allowing for generalization beyond specific objects. A closed-loop physics verifier ensures the validity of generations by applying kinematic feasibility and collision constraints. This resource-efficient framework aims to enhance robot learning from human videos.
Key facts
- Pegasus is a low-resource framework for embodied AI.
- It bridges the embodiment gap between human morphology and robot hardware.
- It uses a graph-based intermediate representation: Task Graph, Affordance Graph, Constraint Graph, Robot Planning Graph.
- A hierarchical affordance latent space enables generalization beyond object identities.
- A closed-loop physics verifier filters invalid generations.
- The framework translates human demonstrations into robot-learnable data.
- The bottleneck in embodied AI is data, not model architecture.
- The preprint is on arXiv with ID 2607.26903.
Entities
Institutions
- arXiv