New AI Task and Benchmark for Person-Centric Video Reasoning Introduced
To tackle the shortcomings in video reasoning, researchers have unveiled the Identity-conditioned Queries (ICQ) task, which challenges the conventional video-text framework that often oversimplifies identity matching and person-focused reasoning. This task compels models to connect and analyze both an input video and a corresponding reference image of an individual, facilitating identity grounding, behavior comprehension, and temporal reasoning. To aid this endeavor, they introduce ISYV (I Seek You in Videos), comprising three key elements: ISYV-Bench, a rigorous evaluation benchmark featuring 1,377 complex real-world videos and 1,377 question-answer pairs across six difficulty levels; ISYV-75K, a comprehensive training dataset with 75K high-quality samples; and a model solution. This research is documented in a paper available on arXiv (arXiv:2608.07417).
Key facts
- ICQ task introduced for joint video and reference image reasoning
- ISYV-Bench includes 1,377 videos and 1,377 QA pairs
- Six difficulty levels from identity recognition to causal reasoning
- ISYV-75K training set contains 75K samples
- Paper available on arXiv with ID 2608.07417
- Addresses limitations of existing video reasoning tasks
- Focus on person-centric reasoning and identity matching
Entities
Institutions
- arXiv