6 papers · 1 filter
Evaluating Compositional Structure in Audio Representations
Chuyang Chen, Bea Steers, Brian McFee +1
We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attr…
Latent Multi-view Learning for Robust Environmental Sound Representations
Sivan Ding, Julia Wilkins, Magdalena Fuentes +1
Self-supervised learning (SSL) approaches, such as contrastive and generative methods, have advanced environmental sound representation learning using unlabeled data. However, how…
Controllable Embedding Transformation for Mood-Guided Music Retrieval
Julia Wilkins, Jaehun Kim, Matthew E. P. Davies +2
Music representations are the backbone of modern recommendation systems, powering playlist generation, similarity search, and personalized discovery. Yet most embeddings offer litt…
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
Julia Wilkins, Sivan Ding, Magdalena Fuentes +1
Recent advances in self-supervised learning (SSL) methods offer a range of strategies for capturing useful representations from music audio without the need for labeled data. While…
Latent Acoustic Mapping for Direction of Arrival Estimation: A Self-Supervised Approach
Adrian S. Roman, Iran R. Roman, Juan P. Bello
Acoustic mapping techniques have long been used in spatial audio processing for direction of arrival estimation (DoAE). Traditional beamforming methods for acoustic mapping, while…
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
Julia Wilkins, Sivan Ding, Magdalena Fuentes +1
Self-supervised learning (SSL) offers a powerful way to learn robust, generalizable representations without labeled data. In music, where labeled data is scarce, existing SSL metho…