23 citations · 39 across the 6 of their papers we have counts for
10 papers
Evaluating Off-the-Shelf Machine Listening and Natural Language Models for Automated Audio Captioning
Benno Weck, Xavier Favory, Konstantinos Drossos +1
Automated audio captioning (AAC) is the task of automatically generating textual descriptions for general audio signals. A captioning system has to identify various information fro…
Enriched Music Representations with Multiple Cross-modal Contrastive Learning
Andres Ferraro, Xavier Favory, Konstantinos Drossos +2
Modeling various aspects that make a music piece unique is a challenging task, requiring the combination of multiple sources of information. Deep learning is commonly used to obtai…
Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and Tags
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen +1
Self-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Publish…
COALA: Co-Aligned Autoencoders for Learning Semantically Enriched Audio Representations
Xavier Favory, Konstantinos Drossos, Tuomas Virtanen +1
Audio representation learning based on deep neural networks (DNNs) emerged as an alternative approach to hand-crafted features. For achieving high performance, DNNs often need a la…
Search Result Clustering in Collaborative Sound Collections
Xavier Favory, Frederic Font, Xavier Serra
The large size of nowadays' online multimedia databases makes retrieving their content a difficult and time-consuming task. Users of online sound collections typically submit searc…
Neural Percussive Synthesis Parameterised by High-Level Timbral Features
António Ramires, Pritish Chandna, Xavier Favory +2
We present a deep neural network-based methodology for synthesising percussive sounds with control over high-level timbral characteristics of the sounds. This approach allows for i…