CrossProjection: New Benchmark Tests AI Models on Architectural Drawing Comprehension
A new research paper introduces CrossProjection, a benchmark designed to evaluate how vision-language models (VLMs) understand architectural drawings. The study, available on arXiv (2608.00473), addresses a fundamental challenge: architectural drawings—plans, sections, and elevations—represent different types of projections (cuts vs. facade projections) that violate assumptions of multi-view reasoning. CrossProjection assesses models on three tasks: Matching, Registration, and Geometric Grounding, using categorical judgments, candidate selection, and free-form localization of points, lines, and regions. The benchmark was tested on 23 real drawing sets with 1,954 categorical conditions per model. Results show GPT-5.5 achieving 82.4% accuracy, Qwen3-VL-32B-Instruct scoring 62.2%, and GLM-4.5V reaching 57.2%. A matched 200-target study compared natural and vector-text-suppressed drawings with closed-candidate and free-geometry outputs. While candidate-supported performance was often higher, free localization remained fragile, particularly on natural drawings. The paper highlights the difficulty VLMs face in preserving component identity and externalizing geometry across heterogeneous architectural views, suggesting that current models lack robust geometric grounding in architectural contexts. The research is significant for the intersection of AI and architecture, as it provides a diagnostic tool for improving model capabilities in understanding spatial representations.
Key facts
- CrossProjection is a benchmark for evaluating vision-language models on architectural drawings.
- It tests Matching, Registration, and Geometric Grounding.
- The study used 23 real drawing sets and 1,954 categorical conditions per model.
- GPT-5.5 scored 82.4%, Qwen3-VL-32B-Instruct 62.2%, and GLM-4.5V 57.2%.
- A matched 200-target study compared natural and vector-text-suppressed drawings.
- Free localization remained fragile, especially on natural drawings.
- Architectural drawings violate assumptions of multi-view reasoning.
- The paper is available on arXiv (2608.00473).
Entities
Institutions
- arXiv