CARE Framework Enhances Medical VQA Accuracy and Calibration
A new framework called CARE (Confidence-Aware Reasoning) aims to improve the reliability of medical visual question answering (VQA) systems by addressing confidence miscalibration in multimodal large language models (MLLMs). The framework, detailed in a paper on arXiv (2608.10964), combines reinforcement fine-tuning (RFT) with a dual-stage pipeline. First, it uses scalable Medical-CoT synthesis to generate structured cold-start data for supervised fine-tuning. Second, it employs Group Relative Policy Optimization (GRPO) with a novel Confidence-Aware Reward (CAR) mechanism that ties the model's confidence to diagnostic correctness. Across three medical VQA benchmarks, CARE achieves the highest diagnostic accuracy and the lowest calibration error, indicating more trustworthy AI-assisted diagnostics. The work is significant for clinical applications where accurate confidence estimates are crucial for decision-making.
Key facts
- CARE stands for Confidence-Aware medical REasoning framework.
- It addresses confidence miscalibration in medical MLLMs.
- Uses dual-stage pipeline: Medical-CoT synthesis for SFT, then GRPO with CAR.
- CAR ties model confidence to diagnostic correctness in reward signal.
- Evaluated on three Medical VQA benchmarks.
- Achieves highest diagnostic accuracy and lowest calibration error.
- Paper available on arXiv with ID 2608.10964.
- Aims to improve clinical trust in AI-assisted diagnostics.
Entities
Institutions
- arXiv