6 papers
Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification
Emiliano Acevedo, Martín Rocamora, Magdalena Fuentes
Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely dr…
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
Julia Wilkins, Sivan Ding, Magdalena Fuentes +1
Self-supervised learning (SSL) offers a powerful way to learn robust, generalizable representations without labeled data. In music, where labeled data is scarce, existing SSL metho…
Two vs. Four-Channel Sound Event Localization and Detection
Julia Wilkins, Magdalena Fuentes, Luca Bondi +3
Sound event localization and detection (SELD) systems estimate both the direction-of-arrival (DOA) and class of sound sources over time. In the DCASE 2022 SELD Challenge (Task 3),…
Sound Source Distance Estimation in Diverse and Dynamic Acoustic Conditions
Saksham Singh Kushwaha, Iran R. Roman, Magdalena Fuentes +1
Localizing a moving sound source in the real world involves determining its direction-of-arrival (DOA) and distance relative to a microphone. Advancements in DOA estimation have be…
Bridging High-Quality Audio and Video via Language for Sound Effects Retrieval from Visual Queries
Julia Wilkins, Justin Salamon, Magdalena Fuentes +2
Finding the right sound effects (SFX) to match moments in a video is a difficult and time-consuming task, and relies heavily on the quality and completeness of text metadata. Retri…
Adapting Meter Tracking Models to Latin American Music
Lucas S. Maia, Martín Rocamora, Luiz W. P. Biscainho +1
Beat and downbeat tracking models have improved significantly in recent years with the introduction of deep learning methods. However, despite these improvements, several challenge…