NewEvery arXiv paper, its researchers & institutions — mapped.
audio processing

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

arXiv:2607.27109

summary

The paper introduces MMAC, a large benchmark of 5,638 audio clips designed to evaluate audio captioning models across multiple capability categories and evaluation dimensions, focusing on information coverage and description reliability.

Abstract

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.

Topics & keywords

#audio captioning#benchmark#evaluation metrics#large language models#multidimensional assessmentAudioLLMMMACinformation coveragedescription reliabilityevaluation dimensionsaudio dataset