Diagram-MMU: New Benchmark Tests AI's Scientific Diagram Parsing
A new benchmark called Diagram-MMU has been launched by researchers to assess Multimodal Large Language Models (MLLMs) in the realm of scientific diagram interpretation and analysis. This benchmark includes 3.7k carefully selected diagrams and 18.3k questions validated by humans across six distinct fields. It evaluates models through three specific tasks: parsing diagrams into code, editing code derived from diagrams, and answering questions related to diagrams, all within agentic contexts. An analysis of 12 MLLMs indicated that tasks involving diagram-to-code are more complex than answering diagram-related questions, suggesting that while models can analyze diagrams, they find parsing and editing more difficult. The benchmark arises from the increasing application of MLLMs in scientific documentation, exemplified by OpenAI Prism, which translates diagrams into LaTeX TikZ code. These results highlight the necessity for enhanced diagram-to-code functionalities.
Key facts
- Diagram-MMU is a new multi-modal benchmark for scientific diagrams.
- It includes 3.7k curated diagrams and 18.3k human-validated questions.
- The benchmark covers six domains.
- It evaluates three tasks: diagram-to-code parsing, diagram-to-code editing, and diagram question answering.
- 12 MLLMs were evaluated.
- Diagram-to-code tasks are more challenging than diagram question answering.
- OpenAI Prism is mentioned as a workspace for scientific writing.
- The paper is available on arXiv (2608.12262).
Entities
Institutions
- OpenAI
- arXiv