ARTFEED — Contemporary Art Intelligence

PinpointQA: Benchmark for Small Object Spatial Understanding in Indoor Videos

ai-technology · 2026-08-13

A new assessment framework named PinpointQA has been launched to evaluate the capability of multimodal large language models (MLLMs) in identifying small, everyday objects in indoor video footage. Utilizing the ScanNet++ and ScanNet200 datasets, which include over 1,000 scenes and more than 10,000 question-answer combinations, the tool provides four levels of increasing difficulty. These include tasks related to object verification, reference identification, detailed spatial descriptions, and structured spatial predictions. PinpointQA fills a crucial gap in current evaluations of video spatial intelligence and object localization, as detailed in an accompanying paper on arXiv, identifier 2604.08991.

Key facts

  • PinpointQA is a benchmark for small object-centric spatial understanding in indoor videos.
  • It is built from ScanNet++ and ScanNet200 datasets.
  • The benchmark contains 1,024 scenes and 10,094 QA pairs.
  • It includes four tasks: Target Presence Verification, Nearest Reference Identification, Fine-Grained Spatial Description, and Structured Spatial Prediction.
  • Ground-truth annotations are constructed from intermediate spatial representations derived from aligned 3D geometry.
  • The benchmark evaluates multimodal large language models (MLLMs).
  • It addresses the lack of direct evaluation of precise localization of small objects in existing benchmarks.
  • The paper is available on arXiv with ID 2604.08991.

Entities

Institutions

  • arXiv

Sources