ARTFEED — Contemporary Art Intelligence

PhysMind: Training-Free Physical Reasoning from Video

ai-technology · 2026-08-06

A new framework called PhysMind has been developed by researchers to enhance physical reasoning in vision-language models (VLMs) without requiring training. This agentic system generates reusable, question-agnostic executable worlds from videos by recovering dynamic scenes through object segmentation, mesh reconstruction, and 6D pose tracking. It then fits continuous-time dynamics and latent physical parameters without the need for a time-stepped simulator. When posed with a question, PhysMind can inspect, edit, or continue the world, providing answers based on the resulting interactions. Compared to traditional chain-of-thought (CoT) reasoning using the same VLM, PhysMind shows a significant accuracy increase of 38.23 points on CLEVRER and 8.08 points on Physion++. The paper can be found on arXiv with the identifier 2608.04575.

Key facts

  • PhysMind is a training-free agentic framework for physical reasoning from video.
  • It constructs one reusable, question-agnostic executable world per video.
  • The framework uses object segmentation, mesh reconstruction, and 6D pose tracking.
  • It fits analytic continuous-time dynamics and latent physical parameters without time-stepped simulation.
  • PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++.
  • The paper is available on arXiv with identifier 2608.04575.
  • The framework is designed to work with vision-language models (VLMs).
  • It answers questions by inspecting, continuing, or editing the executable world.

Entities

Institutions

  • arXiv

Sources