SiPhy: Single-Image Physical Property Reasoning Framework
A new framework named SiPhy has been developed by researchers to deduce physical characteristics like mass, stiffness, and elasticity from a single RGB image. Unlike traditional methods that depend on multi-view reconstruction or physics-based guidance, SiPhy integrates 3D-aware visual information and depth with material knowledge derived from language models. It generates pseudo-voxel points, retrieves CLIP features, and aligns them with material options suggested by a vision-language model (VLM). A part-based contrastive aggregator ensures consistency across regions, while a heaviness-aware refinement enhances the accuracy of thickness and volume estimates for dense objects. SiPhy sets a new benchmark in single-image performance on ABO-500, MVImgNet-100, and PhysXNet-100, outperforming multi-view methods by improving mass MnRE by up to 93% (compared to PUGS) and decreasing density MAE by 35.5% (against NeRF2Physics).
Key facts
- SiPhy infers mass, stiffness, and elasticity from a single image.
- It uses CLIP features and a VLM for material grounding.
- Part-based contrastive aggregator ensures region consistency.
- Heaviness-aware refinement improves thickness and volume estimation.
- Tested on ABO-500, MVImgNet-100, and PhysXNet-100 datasets.
- Outperforms multi-view methods: mass MnRE improved by 93% vs. PUGS.
- Density MAE reduced by 35.5% vs. NeRF2Physics.
- Published on arXiv with ID 2607.22355.
Entities
Institutions
- arXiv