ARTFEED — Contemporary Art Intelligence

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

ai-technology · 2026-07-30

The introduction of a new standard known as MMAC (Massive Multi-dimensional benchmark for Audio Captioning) aims to assess audio captioning capabilities in large language models. This benchmark comprises 5,638 audio clips sourced from more than 20 different datasets, spanning 6 capability categories and 15 evaluation dimensions. It evaluates if the captions produced by models include pertinent information and if this information aligns with reference labels. Analyses of both open-source and proprietary AudioLLMs highlighted significant variations in dimensions, information coverage, and the reliability of descriptions. Further details can be found in arXiv:2607.27109.

Key facts

  • MMAC is a benchmark for audio captioning.
  • It contains 5,638 audio clips from over 20 data sources.
  • It covers 6 capability categories and 15 evaluation dimensions.
  • It checks information coverage and description reliability.
  • Evaluated open-source and proprietary AudioLLMs.
  • Results show differences across evaluation dimensions.
  • Paper available on arXiv:2607.27109.
  • Focuses on moving from brief to open-ended descriptions.

Entities

Institutions

  • arXiv

Sources