EcoFrame: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
A novel framework, EcoFrame, has been introduced to enhance the efficiency of long-video comprehension in vision-language models (VLMs). This method, outlined in a paper on arXiv (2608.03918), tackles the difficulty of selecting sparse visual evidence from lengthy videos. Current techniques either employ static one-shot selection with predetermined frame limits or depend on agent-based schedulers that necessitate expensive multi-round reasoning. EcoFrame features a query-adaptive scheduling system that utilizes feedback from the VLM’s inference to decide when to augment the frame budget and identify additional candidate evidence. It incorporates entropy-gated budget scheduling, which halts early if the existing evidence suffices or expands the budget progressively otherwise. Furthermore, attention-guided candidate proposals transform frame-level attention into a temporal prior, facilitating dense local searches in key areas. The framework is intentionally low-overhead and does not require any training. The paper was released as a cross-type submission on arXiv.
Key facts
- EcoFrame is a training-free framework for efficient long-video understanding.
- It uses entropy-gated budget scheduling to adaptively adjust frame budgets.
- Attention-guided candidate proposal converts frame-level attention into a temporal prior.
- The method leverages VLM inference feedback for query-adaptive scheduling.
- It aims to reduce overhead compared to agent-based schedulers.
- The paper is available on arXiv with ID 2608.03918.
- The announcement type is 'cross'.
- The framework targets vision-language models (VLMs).
Entities
Institutions
- arXiv