4 papers · 1 filter
Woosh: A Sound Effects Foundation Model
Gaëtan Hadjeres, Marc Ferras, Khaled Koutini +7
The audio research community depends on open generative models as foundational tools for building novel approaches and establishing baselines. In this report, we present Woosh, Son…
CrossMuSim: A Cross-Modal Framework for Music Similarity Retrieval with LLM-Powered Text Description Sourcing and Mining
Tristan Tsoi, Jiajun Deng, Yaolong Ju +3
Music similarity retrieval is fundamental for managing and exploring relevant content from large collections in streaming platforms. This paper presents a novel cross-modal contras…
The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?
Pedro Ramoneda, Emilia Parada-Cabaleiro, Benno Weck +1
In this work, we explore the use and reliability of Large Language Models (LLMs) in musicology. From a discussion with experts and students, we assess the current acceptance and co…
MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models
Benno Weck, Ilaria Manco, Emmanouil Benetos +3
Multimodal models that jointly process audio and language hold great promise in audio understanding and are increasingly being adopted in the music domain. By allowing users to que…