VDC-Agent: Self-Evolving Video Captioning via Agentic Self-Reflection
A new autonomous framework for Video Detailed Captioning (VDC) called VDC-Agent has been unveiled by researchers. This system allows a single Multimodal Large Language Model (MLLM) to create and enhance high-quality captions through principle-guided self-reflection, thus removing the necessity for expensive human annotations or proprietary model distillation. To tackle the issue of inference latency during iterative refinement, the model incorporates reflective capabilities internally. The team developed VDC-Agent-19K, a preference dataset based on the agent's self-evaluated trajectories, and introduced a Curriculum Direct Preference Optimization (DPO) method that gradually aligns the model from simpler to more complex samples. Extensive testing shows that VDC-Agent achieves top-tier performance, as detailed in the paper arXiv:2511.19436v2, marking a significant advancement in self-supervised video understanding and multimodal AI.
Key facts
- VDC-Agent is an autonomous self-evolving framework for Video Detailed Captioning.
- It uses a single Multimodal Large Language Model (MLLM) to generate and refine captions.
- The framework relies on principle-guided self-reflection, not human annotations or proprietary models.
- It internalizes reflective capability to overcome inference latency.
- VDC-Agent-19K is a preference dataset derived from the agent's self-scored trajectories.
- Curriculum Direct Preference Optimization (DPO) strategy aligns the model from easy to hard samples.
- Extensive experiments show state-of-the-art performance.
- The paper is available on arXiv with ID 2511.19436v2.
Entities
Institutions
- arXiv