1 paper
Lucia Cascone, Valeria Fraenza, Michele Nappi +2
Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input…