OmniPhys: Benchmarking Physical Commonsense in Text-to-Image Generation
A team of researchers has unveiled OmniPhys, a benchmark consisting of 1,551 samples based on a Physical Knowledge Graph aimed at assessing physical commonsense in text-to-image models. This benchmark employs PhET simulations that correspond to standard curricula to develop diagnostic stress tests through a dual-path verification method. Furthermore, they introduce OmniPrompt, an iterative framework that approaches physical alignment as a discrete optimization challenge, combining prompts to mitigate gradient hallucinations triggered by temporary visual artifacts. This research tackles the common disregard for essential physical laws by text-to-image models, even though they exhibit high visual accuracy.
Key facts
- OmniPhys is a benchmark of 1,551 samples grounded in a Physical Knowledge Graph.
- It aligns PhET simulations with standard curricula.
- It uses a dual-path verification protocol for diagnostic stress tests.
- OmniPrompt is an iterative framework for physical alignment.
- OmniPrompt treats physical alignment as a discrete optimization problem.
- The framework aggregates prompts to overcome gradient hallucinations.
- Text-to-image models often violate physical commonsense.
- Existing benchmarks rely on coarse-grained descriptions.
Entities
—