SceneActBench Tests VLM Agents on 3D Scene Actions
SceneActBench has been introduced by researchers as a benchmark aimed at assessing the actions of vision-language model (VLM) agents in complete multi-object 3D environments, rather than just describing or managing individual objects. This benchmark features five unique 3D tasks derived from 210 source instances, resulting in 520 task cases with corresponding input conditions. Each task functions within a cohesive agent-environment loop, where agents are provided with PNG images or sampled video frames, and, when relevant, 3D assets, to interact with the 3D setting. The final results are measured against hidden ground truth using specific geometric metrics. Scores across eleven proprietary VLM configurations varied from 38.6 to 50.2, with no single configuration excelling in all tasks. This research underscores deficiencies in existing evaluation techniques and offers a framework for a more thorough assessment of agent capabilities in 3D spaces. The study can be found on arXiv with the identifier 2607.22393.
Key facts
- SceneActBench is a benchmark for VLM agents acting on 3D scenes.
- It includes five 3D tasks from 210 source instances.
- 520 task cases are generated with paired input conditions.
- Agents receive PNG images or video frames and optional 3D assets.
- Evaluation uses task-specific geometric metrics against hidden ground truth.
- Eleven proprietary VLM configurations were tested.
- Overall scores range from 38.6 to 50.2.
- No single VLM configuration performs consistently across tasks.
Entities
Institutions
- arXiv