ARTFEED — Contemporary Art Intelligence

ChartAnno: New Benchmark Evaluates MLLMs on Chart Annotation

ai-technology · 2026-08-06

ChartAnno has been unveiled by researchers as a benchmark to assess multimodal large language models (MLLMs) specifically for generating chart annotations. This benchmark features 1,200 authentic charts, each accompanied by code and annotation instructions that vary in specificity across three levels. The evaluation involved ten representative MLLMs tested under two main input conditions: using chart code alone and with both chart code and chart images, alongside a separate study focusing solely on chart images. Findings reveal that proprietary models generally outperform their counterparts, although large-scale open-source models are closing the gap. This study emphasizes chart annotation, a complex communicative task that necessitates models to interpret intended messages and chart semantics. The full research is accessible on arXiv (arXiv:2608.03464).

Key facts

  • ChartAnno is a new benchmark for evaluating MLLMs on chart annotation generation.
  • It contains 1,200 real-world charts with paired code and annotation instructions.
  • Instructions are provided at three levels of specificity.
  • Ten representative MLLMs were evaluated.
  • Two primary input settings were tested: chart code alone, and chart code plus chart image.
  • A chart image-only ablation study was also conducted.
  • Proprietary models are stronger overall, but open-source models are closing the gap.
  • The paper is available on arXiv with ID 2608.03464.

Entities

Institutions

  • arXiv

Sources