4 papers
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
Å imon SedláÄek, Sara Barahona, Bolaji Yusuf +9
Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks rapidly evolve to incorporate complex reas…
Robustness assessment of large audio language models in multiple-choice evaluation
Fernando López, Santosh Kesiraju, Jordi Luque
Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. However, subtle changes, such as sh…
Automatic detection of CMEs using synthetically-trained Mask R-CNN
Francisco A. Iglesias, Diego G. Lloveras, Florencia L. Cisterna +6
Coronal mass ejections (CMEs) are a major driver of space weather. To assess CME geoeffectiveness, among other scientific goals, it is necessary to reliably identify and characteri…
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
Sonal Kumar, Å imon SedláÄek, Vaibhavi Lokegaonkar +31
Audio comprehension-including speech, non-speech sounds, and music-is essential for achieving human-level intelligence. Consequently, AI agents must demonstrate holistic audio unde…