PolyComp Benchmark Tests AI Spatial Reasoning with Polycube Puzzles
Researchers have introduced PolyComp, a procedurally generated benchmark designed to stress visual recognition and compositional spatial reasoning in multimodal AI models. The benchmark presents models with a target solid and four options, requiring identification of which pair of polycube components can combine to form the target. It comprises 120 problems across four geometry families, each presented in three formats using single or multiple images, totaling 360 presented problems per model. The random guessing baseline is 25%. In evaluations, GPT-5.6 Sol with max effort achieved 50.0% accuracy (95% CI 43.3-56.7%) at a mean cost of $0.951 per presented problem. Claude Fable 5 with max effort scored 39.4% (33.1-46.1%) at $0.701, while Gemini 3.1 Pro Preview with thinking level high reached 27.5% (22.8-32.5%), near the random baseline, at $0.350. The accuracy spread across geometry families indicates varying difficulty. The benchmark is released as an arXiv preprint (arXiv:2608.14741) and is intended to advance research in compositional spatial reasoning for AI systems.
Key facts
- PolyComp is a procedurally generated benchmark for compositional 3D spatial reasoning.
- It contains 120 problems across four geometry families.
- Each problem has three presentation formats using single or multiple images.
- Random guessing baseline is 25%.
- GPT-5.6 Sol achieved 50.0% accuracy at $0.951 per problem.
- Claude Fable 5 achieved 39.4% accuracy at $0.701 per problem.
- Gemini 3.1 Pro Preview achieved 27.5% accuracy at $0.350 per problem.
- The benchmark is available on arXiv with ID 2608.14741.
Entities
Institutions
- arXiv