MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
The introduction of a new standard known as MMAC (Massive Multi-dimensional benchmark for Audio Captioning) aims to assess audio captioning capabilities in large language models. This benchmark comprises 5,638 audio clips sourced from more than 20 different datasets, spanning 6 capability categories and 15 evaluation dimensions. It evaluates if the captions produced by models include pertinent information and if this information aligns with reference labels. Analyses of both open-source and proprietary AudioLLMs highlighted significant variations in dimensions, information coverage, and the reliability of descriptions. Further details can be found in arXiv:2607.27109.
Key facts
- MMAC is a benchmark for audio captioning.
- It contains 5,638 audio clips from over 20 data sources.
- It covers 6 capability categories and 15 evaluation dimensions.
- It checks information coverage and description reliability.
- Evaluated open-source and proprietary AudioLLMs.
- Results show differences across evaluation dimensions.
- Paper available on arXiv:2607.27109.
- Focuses on moving from brief to open-ended descriptions.
Entities
Institutions
- arXiv