HSSBench: New Benchmark for Multimodal LLMs in Humanities and Social Sciences
A new benchmark called HSSBench has been developed by researchers to assess Multimodal Large Language Models (MLLMs) specifically in the areas of Humanities and Social Sciences (HSS). This benchmark fills a void in existing evaluation frameworks that mainly emphasize general knowledge and STEM-related reasoning. HSS tasks demand interdisciplinary thinking and the ability to integrate knowledge from various domains, presenting distinct challenges for MLLMs, especially in connecting abstract ideas with visual elements. HSSBench evaluates performance in several languages, including all six official United Nations languages. Additionally, the research introduces an innovative data generation pipeline designed for HSS contexts. The study can be found on arXiv with the identifier 2506.03922.
Key facts
- HSSBench is a new benchmark for evaluating MLLMs on Humanities and Social Sciences tasks.
- Current benchmarks for MLLMs primarily emphasize general knowledge and STEM reasoning.
- HSS tasks require horizontal, interdisciplinary thinking and deep integration of knowledge.
- HSSBench supports multiple languages, including the six official languages of the United Nations.
- A novel data generation pipeline tailored for HSS scenarios is introduced.
- The paper is available on arXiv with identifier 2506.03922.
- The benchmark aims to address the gap in evaluating MLLMs on HSS tasks.
- The work was announced as a replace-cross type on arXiv.
Entities
Institutions
- arXiv