ARTFEED — Contemporary Art Intelligence

ClinMM-Bench: Benchmarking Multi-Turn Multimodal Clinical Reasoning in MLLMs

other · 2026-07-29

ClinMM-Bench has been created by researchers as the most extensive benchmark for multi-turn multimodal clinical diagnostic evaluations, aimed at testing large language models in practical clinical reasoning scenarios. This benchmark features 1,089 complex real-world clinical cases along with 3,760 medical images spanning eight different specialties. A dual-level evaluation framework was employed to assess both the accuracy of diagnoses and the quality of reasoning in 15 selected multimodal large language models (MLLMs). Findings indicated that these models face difficulties with ongoing information disclosure and the need for dynamic updates to hypotheses, revealing shortcomings in existing single-turn evaluation methods.

Key facts

  • ClinMM-Bench is the largest multi-turn multimodal clinical diagnostic evaluation benchmark.
  • It contains 1,089 real-world clinical cases and 3,760 medical images across eight specialties.
  • 15 representative MLLMs were systematically evaluated.
  • A two-level evaluation framework assessed diagnostic accuracy and reasoning quality.
  • The study addresses limitations of single-turn evaluations in capturing clinical complexity.

Entities

Sources