OPD-V: Visual On-Policy Self-Distillation for Multimodal LLMs
A new research paper introduces OPD-V, a visual on-policy self-distillation paradigm designed to address modality imbalance in multimodal large language models (MLLMs). The paper, available on arXiv (2608.05131), argues that existing on-policy self-distillation (OPSD) methods overlook modality imbalance, where textual information dominates generation, preventing the model from fully integrating multimodal input. To study this, the authors construct a Positive Teacher with a Zoom-In Image and a Negative Teacher with a Mask Image, which exhibit different degrees of modality imbalance. Their experiments reveal that modality balance itself can serve as privileged information. Motivated by this, OPD-V instantiates a visual OPSD approach that leverages this insight to improve visual reasoning in MLLMs. The paper is a cross-type announcement, indicating it has been submitted to a conference or journal. The research contributes to the field of AI and machine learning, specifically in improving multimodal reasoning capabilities.
Key facts
- Paper introduces OPD-V, a visual on-policy self-distillation paradigm.
- Addresses modality imbalance in multimodal large language models (MLLMs).
- Existing OPSD methods overlook modality imbalance.
- Positive Teacher uses Zoom-In Image; Negative Teacher uses Mask Image.
- Experiments show modality balance can serve as privileged information.
- OPD-V instantiates visual OPSD to improve visual reasoning.
- Paper available on arXiv with ID 2608.05131.
- Announcement type is cross.
Entities
Institutions
- arXiv