PhysMind: Training-Free Physical Reasoning from Video
A new framework called PhysMind has been developed by researchers to enhance physical reasoning in vision-language models (VLMs) without requiring training. This agentic system generates reusable, question-agnostic executable worlds from videos by recovering dynamic scenes through object segmentation, mesh reconstruction, and 6D pose tracking. It then fits continuous-time dynamics and latent physical parameters without the need for a time-stepped simulator. When posed with a question, PhysMind can inspect, edit, or continue the world, providing answers based on the resulting interactions. Compared to traditional chain-of-thought (CoT) reasoning using the same VLM, PhysMind shows a significant accuracy increase of 38.23 points on CLEVRER and 8.08 points on Physion++. The paper can be found on arXiv with the identifier 2608.04575.
Key facts
- PhysMind is a training-free agentic framework for physical reasoning from video.
- It constructs one reusable, question-agnostic executable world per video.
- The framework uses object segmentation, mesh reconstruction, and 6D pose tracking.
- It fits analytic continuous-time dynamics and latent physical parameters without time-stepped simulation.
- PhysMind improves accuracy by 38.23 points on CLEVRER and 8.08 points on Physion++.
- The paper is available on arXiv with identifier 2608.04575.
- The framework is designed to work with vision-language models (VLMs).
- It answers questions by inspecting, continuing, or editing the executable world.
Entities
Institutions
- arXiv