CURV: Curriculum Learning Framework Enhances Chart Understanding in Multimodal LLMs
A new curriculum learning framework named CURV has been developed by researchers to enhance chart question answering (CQA) in multimodal large language models (MLLMs). This framework tackles the difficulty of merging visual understanding with logical reasoning, a challenge that existing models face due to insufficient visual grounded reasoning capabilities. CURV redefines CQA as a process of multi-step visual grounded reasoning, aligning logical reasoning with dynamic visual grounding through focused spatial attention. To facilitate model training, the researchers created CCQA, a three-tiered curriculum dataset that allows for scalable synthetic generation across various chart types and reasoning approaches. Their findings are published in a paper on arXiv (arXiv:2608.02833), showcasing the potential to improve MLLMs' reasoning accuracy regarding visual information.
Key facts
- CURV is a curriculum learning framework for chart question answering (CQA).
- It targets multimodal large language models (MLLMs).
- CURV reformulates CQA as multi-step visual grounded reasoning.
- Each step coordinates logical reasoning with dynamic visual grounding via spatial attention concentration.
- CCQA is a three-level curriculum dataset introduced for training.
- CCQA uses scalable synthetic generation across diverse chart types and reasoning patterns.
- The paper is available on arXiv with ID 2608.02833.
- The framework aims to improve intrinsic visual grounded reasoning in MLLMs.
Entities
Institutions
- arXiv