6 papers
Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges
Geoffroy Peeters, Zafar Rafii, Magdalena Fuentes +4
In this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large par…
Post-Training Quantization for Audio Diffusion Transformers
Tanmay Khandelwal, Magdalena Fuentes
Diffusion Transformers (DiTs) enable high-quality audio synthesis but are often computationally intensive and require substantial storage, which limits their practical deployment.…
Learning from Silence and Noise for Visual Sound Source Localization
Xavier Juanola, Giovana Morais, Magdalena Fuentes +1
Visual sound source localization is a fundamental perception task that aims to detect the location of sounding sources in a video given its audio. Despite recent progress, we ident…
Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification
Emiliano Acevedo, MartÃn Rocamora, Magdalena Fuentes
Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely dr…
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
Liqian Zhang, Magdalena Fuentes
We present SONIQUE, a model for generating background music tailored to video content. Unlike traditional video-to-music generation approaches, which rely heavily on paired audio-v…
A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
Xavier Juanola, Gloria Haro, Magdalena Fuentes
The task of Visual Sound Source Localization (VSSL) involves identifying the location of sound sources in visual scenes, integrating audio-visual data for enhanced scene understand…