AD2-Bench: New Benchmark for Trustworthy Multimodal Reasoning in Complex Urban Scenes
Researchers have unveiled a new benchmark called AD2-Bench, aimed at enhancing the reliability of Multimodal Large Language Models (MLLMs) in intricate urban environments, particularly in challenging conditions. Detailed in a paper on arXiv (2608.10954), this benchmark incorporates a Hierarchical Visual Diagnosis framework that breaks down reasoning into a structured Chain of Evidence (CoE). This approach facilitates a detailed analysis of reasoning errors, highlighting that effective multimodal reasoning relies heavily on precise evidence gathering. The authors adopt a probabilistic perspective on reasoning and pinpoint two main reasons for reasoning failures. Unlike existing benchmarks that focus solely on outcomes, this one evaluates the reasoning process itself. The full paper can be accessed at https://arxiv.org/abs/2608.10954.
Key facts
- AD2-Bench is a new benchmark for evaluating multimodal reasoning in complex urban scenes.
- It introduces a Hierarchical Visual Diagnosis framework with a structured Chain of Evidence (CoE).
- The benchmark addresses the disconnect between perception and reasoning in MLLMs under adverse conditions.
- Existing outcome-oriented benchmarks fail to diagnose failures in the reasoning process.
- The authors propose that robust multimodal reasoning depends on accurate evidence acquisition.
- Reasoning is formulated from a probabilistic viewpoint, identifying two primary causes of reasoning failure.
- The paper is available on arXiv with identifier 2608.10954.
Entities
Institutions
- arXiv