ARTFEED — Contemporary Art Intelligence

CamChoreo: A Benchmark for Temporally Grounded Camera Motion Understanding

ai-technology · 2026-08-13

A recent study presents CamChoreo, a new benchmark designed for understanding camera motion in videos with temporal grounding and compositionality. This research, available on arXiv (2608.10932v1), highlights the shortcomings of current multimodal large language models (MLLMs) that often provide clip-level labels, neglecting the fact that camera movements can vary within a single shot and occur simultaneously. CamChoreo includes 4,229 authentic single-shot clips, each featuring expert-annotated temporal segments and a concise set of 20 direction-aware labels. Almost 50% of these segments exhibit compound camera motion, showcasing multiple movements within one timeframe. The goal of this research is to enhance spatial intelligence and facilitate controllable video generation by allowing models to pinpoint motion-consistent intervals and recognize all active movements.

Key facts

  • CamChoreo is a benchmark for temporally grounded, compositional camera motion understanding.
  • It contains 4,229 real single-shot clips with expert-annotated temporal segments.
  • The annotations use a compact vocabulary of 20 direction-aware labels.
  • Nearly half of the segments contain compound camera motion.
  • The work addresses limitations of clip-level recognition in MLLMs.
  • It aims to improve spatial intelligence and controllable video generation.
  • The paper is published on arXiv with identifier 2608.10932v1.
  • The research is from a cross-institutional collaboration.

Entities

Institutions

  • arXiv

Sources