4 papers · 1 filter
Evaluating Compositional Structure in Audio Representations
Chuyang Chen, Bea Steers, Brian McFee +1
We propose a benchmark for evaluating compositionality in audio representations. Audio compositionality refers to representing sound scenes in terms of constituent sources and attr…
Skip That Beat: Augmenting Meter Tracking Models for Underrepresented Time Signatures
Giovana Morais, Brian McFee, Magdalena Fuentes
Beat and downbeat tracking models are predominantly developed using datasets with music in 4/4 meter, which decreases their generalization to repertories in other time signatures,…
Hybrid Losses for Hierarchical Embedding Learning
Haokun Tian, Stefan Lattner, Brian McFee +1
In traditional supervised learning, the cross-entropy loss treats all incorrect predictions equally, ignoring the relevance or proximity of wrong labels to the correct answer. By l…
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
Junwon Lee, Modan Tailleur, Laurie M. Heller +5
Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene…