HoloCount: New Benchmark for Visual Counting in MLLMs
A new benchmark called HoloCount has been introduced to evaluate visual counting capabilities in Multimodal Large Language Models (MLLMs). The benchmark is structured around a three-level hierarchical taxonomy, designed to assess models across two main categories: Semantic Counting, which focuses on atomic and property-based enumeration, and Analytical Counting, which evaluates logical composition through spatial and set-based reasoning. The motivation behind HoloCount is the persistent issue of numerical hallucinations in MLLMs, where models often fail to provide accurate quantitative answers despite excelling in qualitative scene understanding. Existing counting benchmarks are criticized for focusing only on basic perception in simplified contexts, thereby missing complex failure modes that arise under logical constraints or adversarial conditions. HoloCount aims to be a holistic and diagnostically rich tool to capture these nuanced challenges. The benchmark is detailed in a paper on arXiv (arXiv:2607.06420), with the announcement type 'cross'. The paper is authored by researchers who have not been named in the provided content. The benchmark's development is significant for advancing multimodal intelligence, as visual counting is a fundamental pillar requiring fine-grained grounding and spatial reasoning. The paper likely includes experiments and results, but those details are not present in the given text. The source URL is https://arxiv.org/abs/2607.06420.
Key facts
- HoloCount is a new benchmark for visual counting in MLLMs.
- It is structured around a three-level hierarchical taxonomy.
- It evaluates Semantic Counting (atomic and property-based enumeration) and Analytical Counting (logical composition via spatial and set-based reasoning).
- MLLMs often exhibit numerical hallucinations, which HoloCount aims to address.
- Existing benchmarks focus on basic perception in simplified contexts.
- HoloCount is designed to capture complex failure modes under logical constraints or adversarial conditions.
- The paper is available on arXiv with ID 2607.06420.
- The announcement type is 'cross'.
Entities
Institutions
- arXiv