5 citations · 10 across the 7 of their papers we have counts for
19 papers
Full-band General Audio Synthesis with Score-based Diffusion
Santiago Pascual, Gautam Bhattacharya, Chunghsin Yeh +2
Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental s…
On loss functions and evaluation metrics for music source separation
Enric Gusó, Jordi Pons, Santiago Pascual +1
We investigate which loss functions provide better separations via benchmarking an extensive set of those for music source separation. To that end, we first survey the most represe…
Adversarial Auto-Encoding for Packet Loss Concealment
Santiago Pascual, Joan Serrà, Jordi Pons
Communication technologies like voice over IP operate under constrained real-time conditions, with voice packets being subject to delays and losses from the network. In such cases,…
PixInWav: Residual Steganography for Hiding Pixels in Audio
Margarita Geleta, Cristina Punti, Kevin McGuinness +3
Steganography comprises the mechanics of hiding data in a host media that may be publicly available. While previous works focused on unimodal setups (e.g., hiding images in images,…
On tuning consistent annealed sampling for denoising score matching
Joan Serrà, Santiago Pascual, Jordi Pons
Score-based generative models provide state-of-the-art quality for image and audio synthesis. Sampling from these models is performed iteratively, typically employing a discretized…
On permutation invariant training for speech source separation
Xiaoyu Liu, Jordi Pons
We study permutation invariant training (PIT), which targets at the permutation ambiguity problem for speaker independent source separation models. We extend two state-of-the-art P…