PinpointQA: Benchmark for Small Object Spatial Understanding in Indoor Videos
A new assessment framework named PinpointQA has been launched to evaluate the capability of multimodal large language models (MLLMs) in identifying small, everyday objects in indoor video footage. Utilizing the ScanNet++ and ScanNet200 datasets, which include over 1,000 scenes and more than 10,000 question-answer combinations, the tool provides four levels of increasing difficulty. These include tasks related to object verification, reference identification, detailed spatial descriptions, and structured spatial predictions. PinpointQA fills a crucial gap in current evaluations of video spatial intelligence and object localization, as detailed in an accompanying paper on arXiv, identifier 2604.08991.
Key facts
- PinpointQA is a benchmark for small object-centric spatial understanding in indoor videos.
- It is built from ScanNet++ and ScanNet200 datasets.
- The benchmark contains 1,024 scenes and 10,094 QA pairs.
- It includes four tasks: Target Presence Verification, Nearest Reference Identification, Fine-Grained Spatial Description, and Structured Spatial Prediction.
- Ground-truth annotations are constructed from intermediate spatial representations derived from aligned 3D geometry.
- The benchmark evaluates multimodal large language models (MLLMs).
- It addresses the lack of direct evaluation of precise localization of small objects in existing benchmarks.
- The paper is available on arXiv with ID 2604.08991.
Entities
Institutions
- arXiv