Spatial-IQ Benchmark Deconstructs MLLM Spatial Reasoning Failures
A new hierarchical diagnostic framework named Spatial-IQ has been developed by researchers to analyze spatial reasoning in multimodal large language models (MLLMs). This framework breaks down spatial reasoning into nine distinct perceptual and cognitive sub-tasks, focusing on object counting within stacked 3D structures, which are categorized by the stages of human spatial cognition, with mental rotation included as an additional assessment. The team utilized NVIDIA Isaac Sim to create a dataset comprising approximately 80,000 stacked 3D structures, complete with ground truth for each task. Current benchmarks view MLLMs as black boxes, complicating the identification of whether failures arise from perceptual challenges, like object boundary recognition, or cognitive difficulties, such as occlusion reasoning. Spatial-IQ seeks to clarify these root causes. This research is documented in arXiv:2607.22864v1.
Key facts
- Spatial-IQ is a hierarchical diagnostic framework for MLLM spatial reasoning.
- It decomposes object counting in stacked 3D structures into 9 sub-tasks.
- Sub-tasks are organized by developmental stages of human spatial cognition.
- Mental rotation is included as an additional target probe.
- Dataset of roughly 80,000 stacked 3D structures generated using NVIDIA Isaac Sim.
- Per-task ground truth is provided for each structure.
- Existing benchmarks fail to identify whether failures are perceptual or cognitive.
- Published on arXiv with ID 2607.22864v1.
Entities
Institutions
- arXiv