ARTFEED — Contemporary Art Intelligence

VideoNorms Dataset Benchmarks Cultural Awareness in Video Language Models

ai-technology · 2026-07-30

A new dataset called VideoNorms has been launched by researchers to assess cultural norm awareness in Video Large Language Models (VideoLLMs). This dataset features annotations from well-known TV shows in the US and China, highlighting instances of cultural norm adherence or breaches, supported by (non-)verbal evidence. Initially, a large VideoLLM annotated each item, followed by evaluations from at least three trained monocultural annotators with deep cultural insights, leading to over 3,000 human assessments. The verification process uncovered differences in norm extraction effectiveness between US and Chinese contexts, suggesting caution with fully automated methods for cultures that are not adequately represented in training datasets. Analysis using hierarchical linear modeling of seven open-weight VideoLLMs indicated poorer performance in specific cultural contexts.

Key facts

  • VideoNorms is a dataset for benchmarking cultural awareness in VideoLLMs.
  • Dataset includes annotations from US and Chinese TV shows.
  • Annotations label adherence or violation of cultural norms.
  • Human-AI collaboration framework used: AI annotation followed by human review.
  • At least three trained monocultural annotators per item.
  • Over 3,000 human judgments collected.
  • Disparity found between US and Chinese norm extraction performance.
  • Seven open-weight VideoLLMs analyzed via hierarchical linear modeling.

Entities

Locations

  • United States
  • China

Sources