ARTFEED — Contemporary Art Intelligence

MMLongBench-Doc-V2: Corrected Annotations and Semantics-Aware Evaluation for Long-Document QA

other · 2026-08-06

The MMLongBench-Doc-V2 is an updated iteration of the MMLongBench-Doc benchmark, which initially featured 1,082 questions across 135 PDFs. This revision tackles two primary concerns that distorted the evaluation of performance: the reference metric compared extracted answers that led to incorrect markings due to numerical formatting discrepancies (e.g., 1,358,000 vs. 1358000), and a considerable number of ground-truth annotations were inaccurate, vague, or incomplete, especially for questions that competent systems answered correctly. The updated version rectifies 106 annotations, providing the corresponding page and calculations for each correction. It replaces the string-based metric with a pinned LLM judge for semantic alignment. Additionally, ten questions with incorrect filenames and one duplicate were eliminated, resulting in 1,071 questions across 134 documents. A key contribution is a decision-making procedure regarding when to widen an empty set key. The benchmark is accessible on arXiv (2608.03397).

Key facts

  • MMLongBench-Doc-V2 corrects 106 annotations in the original benchmark.
  • Each corrected annotation includes the page and arithmetic that justify the correction.
  • The string-based metric is replaced with a pinned LLM judge for semantic evaluation.
  • Ten questions with wrong filenames and one duplicated question were removed.
  • The revised benchmark contains 1,071 questions over 134 documents.
  • The original benchmark had 1,082 questions over 135 PDFs.
  • The revision addresses issues with numerical formatting and annotation quality.
  • The most reusable contribution is a decision procedure for widening empty set keys.

Entities

Institutions

  • arXiv

Sources