CAPS Framework Bridges Agentic Policy Gap in Vision-Text Compression
A recent submission to arXiv (2608.08960) presents CAPS, a two-stage framework for Cross-modal Agentic Policy Self-distillation aimed at bridging the agentic policy gap in vision-text compression for multi-step language-model agents. The study reveals that while converting interaction histories into images lowers context costs, it also introduces a capability gap not solely attributable to OCR quality. The authors conducted controlled assessments of history recovery, matched-state decisions, and complete trajectories, finding that visual-history agents display consistent drift in action selection, query formulation, stopping, and evidence usage. CAPS leverages the more robust text-history policy of the same model to guide its visual-history equivalent, with offline trajectory self-distillation enhancing text-policy behavior for visual-history inputs, complemented by online policy self-distillation for further improvement.
Key facts
- Paper arXiv:2608.08960v1 introduces CAPS framework
- CAPS stands for Cross-modal Agentic Policy Self-distillation
- Addresses agentic policy gap in vision-text compression
- Vision-text compression reduces context costs for language-model agents
- Gap cannot be explained by OCR quality alone
- Visual-history agents show systematic drift in action selection, query formulation, stopping, and evidence use
- Offline trajectory self-distillation transfers text-policy behavior to visual inputs
- Online policy self-distillation is part of the two-stage framework
Entities
Institutions
- arXiv