ARTFEED — Contemporary Art Intelligence

Physics-Grounded Fluid Video Generation with Simulation Dataset and Dual-Stream Supervision

ai-technology · 2026-07-29

A recent study available on arXiv presents a novel technique for enhancing fluid video generation within diffusion models by integrating physics-based supervision. The researchers note that existing video diffusion models, while visually striking, often depict unrealistic fluid behaviors, such as liquid columns breaking in mid-air or water levels failing to rise during pouring. To tackle this issue, they developed a fluid dataset that includes 1,638 MPM-simulated pouring and sloshing videos, alongside 2,320 real pouring videos sourced from stock footage, plus two separate test sets: a benchmark of 1,515 real videos and an 18-prompt text-to-first-frame generalization benchmark. They introduce a dual-stream image-to-video framework based on a pretrained diffusion-transformer video generator, aiming to instill accurate fluid dynamics through dual-stream optical-flow supervision, merging both simulated and real data. The paper is cataloged on arXiv under ID 2607.25321.

Key facts

  • arXiv paper ID 2607.25321
  • Dataset includes 1,638 MPM-simulated videos and 2,320 real pouring videos
  • Held-out test sets: 1,515-video real benchmark and 18-prompt generalization benchmark
  • Architecture: dual-stream image-to-video on pretrained diffusion-transformer
  • Goal: enforce physics-based fluid dynamics in video generation

Entities

Institutions

  • arXiv

Sources