10 papers · 1 filter
TinyMU: A Compact Audio-Language Model for Music Understanding
Xiquan Li, Aurian Quelennec, Slim Essid
Music understanding and reasoning are central challenges in the Music Information Research field, with applications ranging from retrieval and recommendation to music agents and vi…
Contrastive Knowledge Distillation for Embedding Refinement in Personalized Speech Enhancement
Thomas Serre, Mathieu Fontaine, Ãric Benhaim +1
Personalized speech enhancement (PSE) has shown convincing results when it comes to extracting a known target voice among interfering ones. The corresponding systems usually incorp…
Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
Antonin Gagnere, Slim Essid, Geoffroy Peeters
Ambiguities in data and problem constraints can lead to diverse, equally plausible outcomes for a machine learning task. In beat and downbeat tracking, for instance, different list…
MATPAC++: Enhanced Masked Latent Prediction for Self-Supervised Audio Representation Learning
Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters +1
Masked latent prediction has emerged as a leading paradigm in self-supervised learning (SSL), especially for general audio and music representation learning. While recent methods h…
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
Clémentine Berger, Roland Badeau, Slim Essid
People often listen to music in noisy environments, seeking to isolate themselves from ambient sounds. Indeed, a music signal can mask some of the noise's frequency components due…
Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning
Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters +1
Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned l…