1 paper · 1 filter
Lucia Cascone, Valeria Fraenza, Michele Nappi +2
Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input…