12 papers
MusTBENCH: Benchmarking and Advancing Temporal Grounding in Music LLMs
Daeyong Kwon, Qiyu Wu, Shinobu Kuriya +6
Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temp…
Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew
Joan SerrÃ, Dipam Goswami, Fabio Morreale +2
Training data attribution (TDA) should enable generative model interpretability and foster a variety of related downstream tasks. Nonetheless, current TDA approaches lack reliabili…
Woosh: A Sound Effects Foundation Model
Gaëtan Hadjeres, Marc Ferras, Khaled Koutini +7
The audio research community depends on open generative models as foundational tools for building novel approaches and establishing baselines. In this report, we present Woosh, Son…
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
Eleonora Mancini, Joan SerrÃ, Paolo Torroni +1
Audio-based lyrics matching can be an appealing alternative to other content-based retrieval approaches, but existing methods often suffer from limited reproducibility and inconsis…
LLM2Fx-Tools: Tool Calling For Music Post-Production
Seungheon Doh, Junghyun Koo, Marco A. MartÃnez-RamÃrez +5
This paper introduces LLM2Fx-Tools, a multimodal tool-calling framework that generates executable sequences of audio effects (Fx-chain) for music post-production. LLM2Fx-Tools uses…
Emergent, not Immanent: A Baradian Reading of Explainable AI
Fabio Morreale, Joan SerrÃ, Yuki Mitsufuji
Explainable AI (XAI) is frequently positioned as a technical problem of revealing the inner workings of an AI model. This position is affected by unexamined onto-epistemological as…