MultivationBench: Benchmarking Multimodal Motivation Reasoning in AI
A novel benchmark, named MultivationBench, has been launched to assess how multimodal large language models reason about motivations in a sequential manner within narrative-driven visual stories. This benchmark is based on well-known psychological theories, including Maslow's hierarchy of needs and Reiss's basic desires, and it challenges models to synthesize accumulated multimodal information to deduce changing motivations. Preliminary findings indicate that every model evaluated faces difficulties in sustaining coherent motivation reasoning throughout sequential contexts, highlighting a significant gap in the social intelligence of current AI systems. This research is available on arXiv with the identifier 2607.26465.
Key facts
- MultivationBench is a benchmark for multimodal sequential motivation reasoning.
- It evaluates multimodal large language models.
- The benchmark uses Maslow's hierarchy and Reiss's basic desires.
- Models must integrate accumulated multimodal context.
- All tested models struggled with consistent motivation reasoning.
- The research is published on arXiv (2607.26465).
- The benchmark addresses a gap in evaluating cumulative behavioral drivers.
- It focuses on story-driven visual narratives.
Entities
Institutions
- arXiv