FilmBench: Cinematic Video Generation Benchmark from Beijing Film Academy
Researchers from the Beijing Film Academy and Hujing Digital Media & Entertainment Group have developed FilmBench, a text-to-video and reference-to-video benchmark grounded in professional Cinematic Language criteria. Unlike existing benchmarks that use web-sourced prompts and generic multimodal models, FilmBench evaluates video generation based on film-academy standards, co-developed with directors and faculty. Prompts are reverse-engineered from film clips, and the evaluation taxonomy includes professional criteria rather than basic visual quality and text alignment. The benchmark aims to assess film-grade craft rather than mere video plausibility.
Key facts
- FilmBench is a text-to-video (T2V) and reference-to-video (R2V) benchmark.
- It is grounded in professional Cinematic Language of the film-academy tradition.
- Co-developed with directors and faculty from Beijing Film Academy and Hujing Digital Media & Entertainment Group.
- Prompts are reverse-engineered from clips of films.
- Evaluation uses professional Cinematic Language criteria instead of generic multimodal models.
- Existing benchmarks use web-sourced prompts or LLM templates and untrained models.
- Current evaluation taxonomies focus on overall visual quality, coarse text alignment, and temporal smoothness.
- FilmBench assesses film-grade craft rather than basic video plausibility.
Entities
Institutions
- Beijing Film Academy
- Hujing Digital Media & Entertainment Group
Locations
- Beijing
- China