O²-CritiCuRL: Curriculum RL for Multimodal Reasoning
A new framework called O²-CritiCuRL addresses flawed intermediate reasoning steps in multimodal large language models. The system uses an iterative offline-online curriculum reinforcement learning approach. In the offline stage, multi-rollout analysis over step-annotated trajectories estimates step-level importance, distilling critical reasoning steps and filtering redundant ones. The online stage employs progressive step-level RL with truncated chains to guide the model. The work is published on arXiv under ID 2607.23700.
Key facts
- O²-CritiCuRL is a curriculum reinforcement learning framework for multimodal reasoning.
- It addresses flawed intermediate steps in multimodal large language models.
- The framework uses an iterative offline-online paradigm.
- Offline stage: multi-rollout analysis over step-annotated trajectories.
- Online stage: progressive step-level reinforcement learning with truncated chains.
- Published on arXiv with ID 2607.23700.
- The approach aims to improve interpretability and reliability.
- It distinguishes critical steps from redundant ones.
Entities
Institutions
- arXiv