EgoMonth: First Month-Level Egocentric Video Benchmark for Long-Term Memory
A team of researchers has launched EgoMonth, the inaugural benchmark for month-level egocentric video comprehension aimed at evaluating the long-term spatiotemporal memory of Multimodal Large Language Models (MLLMs). This benchmark includes more than 300 hours of first-person recordings from 20 individuals, covering a duration of 20 to 120 days, along with 1,443 carefully crafted multiple-choice questions and answers. It features a 14-task evaluation framework based on cognitive principles, categorized into three levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning. The research assesses both open-source and closed-source MLLMs, uncovering notable deficiencies in their long-term memory skills. This initiative addresses a significant shortcoming of current video benchmarks, which often utilize web-sourced videos that lack continuity in spatiotemporal context. The findings are available in a paper on arXiv (arXiv:2608.13113).
Key facts
- EgoMonth is the first month-level egocentric video understanding benchmark.
- It includes over 300 hours of first-person daily-life recordings from 20 participants.
- Recordings span 20 to 120 days.
- The benchmark includes 1,443 human-crafted multiple-choice question-answer pairs.
- Evaluation framework has 14 tasks across three cognitive levels: Schema Consolidation, Episodic Indexing, and Cascading Reasoning.
- State-of-the-art open-source and closed-source MLLMs were evaluated.
- Existing benchmarks rely on web-sourced videos lacking inter-clip spatiotemporal continuity.
- The paper is available on arXiv with ID 2608.13113.
Entities
Institutions
- arXiv