7 citations · 22 across the 11 of their papers we have counts for
16 papers · 1 filter
Full-band General Audio Synthesis with Score-based Diffusion
Santiago Pascual, Gautam Bhattacharya, Chunghsin Yeh +2
Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental s…
Quantitative Evidence on Overlooked Aspects of Enrollment Speaker Embeddings for Target Speaker Separation
Xiaoyu Liu, Xu Li, Joan Serrà
Single channel target speaker separation (TSS) aims at extracting a speaker's voice from a mixture of multiple talkers given an enrollment utterance of that speaker. A typical deep…
On loss functions and evaluation metrics for music source separation
Enric Gusó, Jordi Pons, Santiago Pascual +1
We investigate which loss functions provide better separations via benchmarking an extensive set of those for music source separation. To that end, we first survey the most represe…
Assessing Algorithmic Biases for Musical Version Identification
Furkan Yesiler, Marius Miron, Joan Serrà +1
Version identification (VI) systems now offer accurate and scalable solutions for detecting different renditions of a musical composition, allowing the use of these systems in indu…
Audio-based Musical Version Identification: Elements and Challenges
Furkan Yesiler, Guillaume Doras, Rachel M. Bittner +2
In this article, we aim to provide a review of the key ideas and approaches proposed in 20 years of scientific literature around musical version identification (VI) research and co…
Adversarial Auto-Encoding for Packet Loss Concealment
Santiago Pascual, Joan Serrà, Jordi Pons
Communication technologies like voice over IP operate under constrained real-time conditions, with voice packets being subject to delays and losses from the network. In such cases,…