ARTFEED — Contemporary Art Intelligence

Synthetic Dataset for Robot Spatial Reasoning via Vision-Language Models

ai-technology · 2026-07-29

A new research paper presents a framework aimed at training Vision-Language Models (VLMs) to execute Visual Perspective Taking (VPT), an essential skill for embodied cognition in Human-Robot Interaction (HRI). The authors have developed a synthetic dataset utilizing NVIDIA Omniverse, tailored for supervised learning in spatial reasoning tasks. Each entry in the dataset consists of an RGB image, a corresponding natural language description, and a ground-truth 4x4 transformation matrix that indicates object pose. Initially, the emphasis is on determining Z-axis distance, with future plans to incorporate full 6 Degrees Of Freedom (DOFs) reasoning. This dataset is made publicly accessible to encourage further exploration in embodied AI for interactive human-robot applications.

Key facts

  • Framework trains VLMs for Visual Perspective Taking (VPT)
  • Synthetic dataset generated in NVIDIA Omniverse
  • Each instance includes RGB image, natural language description, and 4x4 transformation matrix
  • Initial focus on inferring Z-axis distance
  • Future extensions target full 6 Degrees Of Freedom (DOFs) reasoning
  • Dataset is publicly available
  • Work is a foundational step toward embodied AI for HRI
  • Published on arXiv under Computer Science > Artificial Intelligence

Entities

Institutions

  • NVIDIA Omniverse
  • arXiv

Sources