ARTFEED — Contemporary Art Intelligence

C3PO Benchmark Reveals Modality Dominance in Multimodal AI

ai-technology · 2026-08-07

A novel benchmark named C3PO has been introduced in an arXiv paper (2608.05381) to assess cross-modal composition and counterfactual performance in omnimodal models. It consists of 3,404 samples that include video, audio, image, and text, aimed at evaluating two key skills: information composition (integrating scattered evidence) and counterfactual conflict (addressing intentional contradictions). Developed through an automated pipeline utilizing 25 logically grounded templates, C3PO incorporates a paired IC/CC structure and a four-tier design to diagnose cross-modal reasoning failures effectively. Findings indicate that humans achieve 88.64% accuracy, while the top model, Gemini-3.1-Pro, only attains 73.17%. Notably, 86-95% of failures arise from modality dominance, where models focus on one modality and overlook conflicting evidence. This research underscores a critical limitation in current multimodal large language models (MLLMs) and serves as a diagnostic tool for the AI research community and multimodal system developers.

Key facts

  • C3PO is a benchmark with 3,404 samples across video, audio, image, and text.
  • It evaluates information composition and counterfactual conflict.
  • Built via a fully automatic pipeline using 25 logically grounded templates.
  • Humans achieve 88.64% accuracy, while the best model (Gemini-3.1-Pro) reaches 73.17%.
  • Open-source models collapse under conflict.
  • 86-95% of failures are due to modality dominance.
  • The paper is available on arXiv with ID 2608.05381.
  • The study uses attention probes to diagnose failures.

Entities

Institutions

  • arXiv

Sources