ARTFEED — Contemporary Art Intelligence

FilmBench: Cinematic Video Generation Benchmark from Beijing Film Academy

ai-technology · 2026-07-30

Researchers from the Beijing Film Academy and Hujing Digital Media & Entertainment Group have developed FilmBench, a text-to-video and reference-to-video benchmark grounded in professional Cinematic Language criteria. Unlike existing benchmarks that use web-sourced prompts and generic multimodal models, FilmBench evaluates video generation based on film-academy standards, co-developed with directors and faculty. Prompts are reverse-engineered from film clips, and the evaluation taxonomy includes professional criteria rather than basic visual quality and text alignment. The benchmark aims to assess film-grade craft rather than mere video plausibility.

Key facts

  • FilmBench is a text-to-video (T2V) and reference-to-video (R2V) benchmark.
  • It is grounded in professional Cinematic Language of the film-academy tradition.
  • Co-developed with directors and faculty from Beijing Film Academy and Hujing Digital Media & Entertainment Group.
  • Prompts are reverse-engineered from clips of films.
  • Evaluation uses professional Cinematic Language criteria instead of generic multimodal models.
  • Existing benchmarks use web-sourced prompts or LLM templates and untrained models.
  • Current evaluation taxonomies focus on overall visual quality, coarse text alignment, and temporal smoothness.
  • FilmBench assesses film-grade craft rather than basic video plausibility.

Entities

Institutions

  • Beijing Film Academy
  • Hujing Digital Media & Entertainment Group

Locations

  • Beijing
  • China

Sources