23 citations · 40 across the 7 of their papers we have counts for
4 papers · 1 filter
Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen +1
Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Publish…
FSD50K: An Open Dataset of Human-Labeled Sound Events
Eduardo Fonseca, Xavier Favory, Jordi Pons +2
Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2M tracks from YouTube videos an…
COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen +1
Audio representation learning based on deep neural networks (DNNs) emerged as an alternative approach to hand-crafted features. For achieving high performance, DNNs often need a la…
Search Result Clustering in Collaborative Sound Collections
Xavier Favory, Frederic Font, Xavier Serra
The large size of nowadays' online multimedia databases makes retrieving their content a difficult and time-consuming task. Users of online sound collections typically submit searc…