Counterfactual Modality Attribution: New Framework for Multimodal LLMs
A new framework called Counterfactual Modality Attribution (CMA) has been developed by researchers to measure how much each modality—image and text—contributes to the predictions of multimodal large language models (MLLMs). This innovation fills a significant void in explainability, as current techniques can highlight key image areas or text tokens but fail to identify which modality primarily influences a model's output. This distinction is crucial, given that a model may arrive at accurate conclusions based on misleading evidence, which could obscure shortcut learning and unsafe reasoning. CMA creates counterfactuals for image-only, text-only, and joint multimodal scenarios using coupled diffusion priors, transforming them into modality attribution scores through a cooperative game-theoretic approach rooted in Shapley values. This framework is the first of its kind for MLLMs. The research paper can be found on arXiv with the identifier 2608.00076, and its abstract notes a cross-listing. This study underscores the increasing significance of explainability in AI systems that aid in critical decision-making by integrating diverse information from various modalities.
Key facts
- CMA is the first framework for quantifying modality-level contributions in MLLMs.
- It uses coupled diffusion priors to generate image-only, text-only, and joint multimodal counterfactuals.
- Modality attribution scores are derived using Shapley values from cooperative game theory.
- The framework addresses the question of which modality drives a prediction.
- It aims to expose shortcut learning and unsafe reasoning in multimodal models.
- The paper is available on arXiv with ID 2608.00076.
- The research is relevant for high-stakes decision-making applications.
- Existing explainability methods only identify influential image regions or text tokens, not modality-level contributions.
Entities
Institutions
- arXiv