activity
20232025
collaborators

6 papers

cs.SD2025

Domain Adaptation Method and Modality Gap Impact in Audio-Text Models for Prototypical Sound Classification

Emiliano Acevedo, Martín Rocamora, Magdalena Fuentes

Audio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data. However, we show that their performance severely dr…

cs.SD2024

Self-Supervised Multi-View Learning for Disentangled Music Audio Representations

Julia Wilkins, Sivan Ding, Magdalena Fuentes +1

Self-supervised learning (SSL) offers a powerful way to learn robust, generalizable representations without labeled data. In music, where labeled data is scarce, existing SSL metho…

cs.SD2023

Two vs. Four-Channel Sound Event Localization and Detection

Julia Wilkins, Magdalena Fuentes, Luca Bondi +3

Sound event localization and detection (SELD) systems estimate both the direction-of-arrival (DOA) and class of sound sources over time. In the DCASE 2022 SELD Challenge (Task 3),…

cs.SD2023

Sound Source Distance Estimation in Diverse and Dynamic Acoustic Conditions

Saksham Singh Kushwaha, Iran R. Roman, Magdalena Fuentes +1

Localizing a moving sound source in the real world involves determining its direction-of-arrival (DOA) and distance relative to a microphone. Advancements in DOA estimation have be…

cs.SD2023

Bridging High-Quality Audio and Video via Language for Sound Effects Retrieval from Visual Queries

Julia Wilkins, Justin Salamon, Magdalena Fuentes +2

Finding the right sound effects (SFX) to match moments in a video is a difficult and time-consuming task, and relies heavily on the quality and completeness of text metadata. Retri…

cs.SD2023

Adapting Meter Tracking Models to Latin American Music

Lucas S. Maia, Martín Rocamora, Luiz W. P. Biscainho +1

Beat and downbeat tracking models have improved significantly in recent years with the introduction of deep learning methods. However, despite these improvements, several challenge…