6 papers
Evaluating Compositional Structure in Audio Representations
Chuyang Chen, Bea Steers, Brian McFee +1
We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attr…
Skip That Beat: Augmenting Meter Tracking Models for Underrepresented Time Signatures
Giovana Morais, Brian McFee, Magdalena Fuentes
Beat and downbeat tracking models are predominantly developed using datasets with music in 4/4 meter, which decreases their generalization to repertories in other time signatures,…
Investigating the Sensitivity of Pre-trained Audio Embeddings to Common Effects
Victor Deng, Changhong Wang, Gael Richard +1
In recent years, foundation models have significantly advanced data-driven systems across various domains. Yet, their underlying properties, especially when functioning as feature…
Hybrid Losses for Hierarchical Embedding Learning
Haokun Tian, Stefan Lattner, Brian McFee +1
In traditional supervised learning, the cross-entropy loss treats all incorrect predictions equally, ignoring the relevance or proximity of wrong labels to the correct answer. By l…
Sound Scene Synthesis at the DCASE 2024 Challenge
Mathieu Lagrange, Junwon Lee, Modan Tailleur +5
This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and d…
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
Junwon Lee, Modan Tailleur, Laurie M. Heller +5
Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene…