ARTFEED — Contemporary Art Intelligence

ReMem: Training-Free Long Video Understanding via Temporal Granularity-Adaptive Keyframe Selection

ai-technology · 2026-07-29

A novel framework named ReMem (Reasoning with Memory) has been introduced for understanding long videos without the need for training. This framework tackles the challenges posed by Multimodal Large Language Models (MLLMs), which have limited context windows that restrict their capability to analyze lengthy videos. ReMem features a dual-level memory-augmented adaptation: at the query level, Memory-Driven Question Parsing utilizes the long-term memory of LLMs to interpret the temporal granularity of questions and identify semantic entities; at the video level, Synergistic Dual-Semantic Frame Alignment leverages intrinsic structural memory to synchronize frames with query semantics, facilitating Structure-Aware Dynamic Frame Routing for frame clustering. Designed for Long Video Question Answering (LongVideoQA), it tailors keyframe selection according to query temporal granularity, addressing the drawbacks of uniform sampling or static query-guided selection. The paper can be accessed on arXiv under ID 2607.24794.

Key facts

  • ReMem is a temporal granularity-adaptive keyframe selection framework for training-free LongVideoQA.
  • It addresses restricted context windows in MLLMs for long video understanding.
  • The framework uses dual-level memory-augmented adaptation: query-level and video-level.
  • Memory-Driven Question Parsing uses LLM long-term memory to decode question temporal granularity.
  • Synergistic Dual-Semantic Frame Alignment aligns frames with query semantics using intrinsic structural memory.
  • Structure-Aware Dynamic Frame Routing clusters frames based on alignment.
  • The paper is published on arXiv with ID 2607.24794.
  • The framework is training-free, meaning no additional fine-tuning is required.

Entities

Institutions

  • arXiv

Sources