ARTFEED — Contemporary Art Intelligence

MMArch: New Benchmark Tests Multimodal AI Reasoning in Architecture

ai-technology · 2026-08-11

A new benchmark called MMArch has been developed by researchers to assess multimodal large language models (MLLMs) regarding their reasoning skills in architectural and civil engineering contexts. This benchmark, outlined in a paper on arXiv (2608.09281), encompasses ten subdomains and is entirely derived from figures in peer-reviewed literature. It includes 1,212 short-answer questions created through a decoupled planner-writer approach, which underwent automated screening, a blind adversarial audit, and expert evaluation. The questions challenge models to recognize pertinent visual data, discern applicable engineering principles, and implement them, steering clear of shortcuts from textual or single-figure hints. In tests involving 18 open-weight and proprietary MLLMs assessed by a panel of domain experts, the top-performing open-source model only reached approximately 30% accuracy, highlighting a notable disparity between AI capabilities and human expertise. This benchmark seeks to advance MLLMs beyond mere drawing recognition or information retrieval, evaluating their ability to synthesize diverse visual information with engineering insights. This research is significant for the architecture and design sectors as it underscores existing limitations and explores future possibilities for AI in engineering and design.

Key facts

  • MMArch is a new benchmark for multimodal reasoning in architecture and civil engineering.
  • It spans ten subdomains and is built from figures in peer-reviewed papers.
  • The benchmark contains 1,212 short-answer items.
  • Items are generated by a decoupled planner-writer pipeline and validated via automated screening, blind adversarial audit, and expert review.
  • The benchmark tests perception of visual evidence, identification of governing principles, and application of those principles.
  • 18 open-weight and proprietary MLLMs were evaluated against a domain-expert panel.
  • The strongest open-source model achieved about 30% accuracy.
  • The paper is available on arXiv with identifier 2608.09281.

Entities

Institutions

  • arXiv

Sources