ARTFEED — Contemporary Art Intelligence

MERIT: A Simple Multi-Key Episodic Memory Retrieval Framework for Ultra-Long Video Understanding

ai-technology · 2026-08-11

A recent paper on arXiv (2608.07663) introduces MERIT (Multi-key Episodic Retrieval with Inference-time Temporal expansion), a novel approach aimed at understanding ultra-long videos. The researchers contend that existing Multi-modal Large Language Models (MLLMs) struggle with videos lasting from hours to days when processed end-to-end. They propose a two-stage method: first, constructing memory without queries, followed by inference based on retrieval. While previous studies have focused on intricate memory construction to establish high-level video relations, MERIT emphasizes high-recall retrieval during memory creation, postponing complex relation composition until inference. This framework employs an episodic multi-key representation for accurate retrieval of detailed memories through a straightforward key-matching process and incorporates neighbor filtering to improve retrieval efficiency. The paper is classified as a cross-type submission and can be accessed at https://arxiv.org/abs/2608.07663.

Key facts

  • Paper arXiv:2608.07663 proposes MERIT framework for ultra-long video understanding.
  • MERIT stands for Multi-key Episodic Retrieval with Inference-time Temporal expansion.
  • The framework addresses videos lasting from hours to days, which are impractical for current MLLMs.
  • It uses a two-stage paradigm: query-agnostic memory construction and retrieval-based inference.
  • MERIT prioritizes high-recall retrievability during memory building.
  • It defers query-specific, high-level relation composition to inference time.
  • The method uses an episodic multi-key representation for precise retrieval.
  • It introduces neighbor filtering to enhance retrieval.

Entities

Institutions

  • arXiv

Sources