Synthetic Dataset for Robot Spatial Reasoning via Vision-Language Models
A new research paper presents a framework aimed at training Vision-Language Models (VLMs) to execute Visual Perspective Taking (VPT), an essential skill for embodied cognition in Human-Robot Interaction (HRI). The authors have developed a synthetic dataset utilizing NVIDIA Omniverse, tailored for supervised learning in spatial reasoning tasks. Each entry in the dataset consists of an RGB image, a corresponding natural language description, and a ground-truth 4x4 transformation matrix that indicates object pose. Initially, the emphasis is on determining Z-axis distance, with future plans to incorporate full 6 Degrees Of Freedom (DOFs) reasoning. This dataset is made publicly accessible to encourage further exploration in embodied AI for interactive human-robot applications.
Key facts
- Framework trains VLMs for Visual Perspective Taking (VPT)
- Synthetic dataset generated in NVIDIA Omniverse
- Each instance includes RGB image, natural language description, and 4x4 transformation matrix
- Initial focus on inferring Z-axis distance
- Future extensions target full 6 Degrees Of Freedom (DOFs) reasoning
- Dataset is publicly available
- Work is a foundational step toward embodied AI for HRI
- Published on arXiv under Computer Science > Artificial Intelligence
Entities
Institutions
- NVIDIA Omniverse
- arXiv