Training-Free and Agentic Reasoning for Video Anomaly Detection
A recent paper published on arXiv (2608.11260) presents a comprehensive reasoning framework for Video Anomaly Detection (VAD), tackling the disconnect between 'when' and 'what' found in current techniques. While traditional DNN methods can identify anomalies in time, they often lack semantic insight. Conversely, LLM-based approaches provide explanations but fail to achieve accurate temporal localization. The authors introduce 'Glance then Scrutinize' (GtS), a framework that operates without training, utilizing both static and dynamic textual cues for effective anomaly grounding and comprehension, optimizing both precision and efficiency. Furthermore, they propose a tool-enhanced agentic VAD strategy to address the shortcomings of static external modules. This research draws inspiration from human analysis of surveillance footage, which involves forming temporal hypotheses through initial glances and then closely examining questionable segments. The paper can be accessed at https://arxiv.org/abs/2608.11260.
Key facts
- Paper arXiv:2608.11260 proposes a unified reasoning paradigm for VAD.
- Existing VAD methods exhibit a 'when-what' dissociation.
- GtS is a training-free framework using static and dynamic textual guidance.
- GtS performs coarse-to-fine anomaly grounding and understanding.
- A tool-augmented agentic VAD approach is proposed to break the ceiling of frozen external modules.
- The paradigm is inspired by human inspection of surveillance videos.
- The paper is available on arXiv.
- The approach balances accuracy and speed.
Entities
Institutions
- arXiv