ARTFEED — Contemporary Art Intelligence

CAPS Framework Bridges Agentic Policy Gap in Vision-Text Compression

ai-technology · 2026-08-11

A recent submission to arXiv (2608.08960) presents CAPS, a two-stage framework for Cross-modal Agentic Policy Self-distillation aimed at bridging the agentic policy gap in vision-text compression for multi-step language-model agents. The study reveals that while converting interaction histories into images lowers context costs, it also introduces a capability gap not solely attributable to OCR quality. The authors conducted controlled assessments of history recovery, matched-state decisions, and complete trajectories, finding that visual-history agents display consistent drift in action selection, query formulation, stopping, and evidence usage. CAPS leverages the more robust text-history policy of the same model to guide its visual-history equivalent, with offline trajectory self-distillation enhancing text-policy behavior for visual-history inputs, complemented by online policy self-distillation for further improvement.

Key facts

  • Paper arXiv:2608.08960v1 introduces CAPS framework
  • CAPS stands for Cross-modal Agentic Policy Self-distillation
  • Addresses agentic policy gap in vision-text compression
  • Vision-text compression reduces context costs for language-model agents
  • Gap cannot be explained by OCR quality alone
  • Visual-history agents show systematic drift in action selection, query formulation, stopping, and evidence use
  • Offline trajectory self-distillation transfers text-policy behavior to visual inputs
  • Online policy self-distillation is part of the two-stage framework

Entities

Institutions

  • arXiv

Sources