ARTFEED — Contemporary Art Intelligence

Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal LLMs via Cross-Frame Visual Prompting

ai-technology · 2026-08-13

The newly introduced framework, Motion-as-Prompt (MaP), seeks to enhance motion-focused video reasoning within multimodal large language models (MLLMs). Conventional MLLMs utilize sparse uniform sampling for video analysis, which can overlook essential transitions between frames, thereby hindering the understanding of object movements, collisions, and causal relationships. MaP addresses this by retrieving dense point trajectories, choosing frames that convey motion information, and overlaying the trajectories accumulated between consecutive sampled frames onto the visual inputs. This approach makes hidden displacements, directional shifts, and interactions visible to static MLLMs. Testing on CLEVRER and Something-Something-v2 demonstrates that MaP significantly boosts average motion-reasoning accuracy. This research can be found on arXiv with the identifier 2608.11655.

Key facts

  • MaP is a track-guided cross-frame visual prompting framework.
  • It recovers dense point trajectories and selects motion-informative frames.
  • It marks trajectories between consecutive sampled frames onto visual inputs.
  • It makes hidden displacement, direction changes, and interactions observable to frozen MLLMs.
  • Experiments on CLEVRER and Something-Something-v2 show consistent improvement in motion-reasoning accuracy.
  • The paper is available on arXiv with ID 2608.11655.
  • The research addresses limitations of sparse uniform sampling in MLLMs.
  • Motion-centric video reasoning is fundamental to robotic manipulation and autonomous navigation.

Entities

Institutions

  • arXiv

Sources