7 citations · 20 across the 6 of their papers we have counts for
21 papers
On tuning consistent annealed sampling for denoising score matching
Joan Serrà, Santiago Pascual, Jordi Pons
Score-based generative models provide state-of-the-art quality for image and audio synthesis. Sampling from these models is performed iteratively, typically employing a discretized…
Investigating the efficacy of music version retrieval systems for setlist identification
Furkan Yesiler, Emilio Molina, Joan Serrà +1
The setlist identification (SLI) task addresses a music recognition use case where the goal is to retrieve the metadata and timestamps for all the tracks played in live music event…
Automatic multitrack mixing with a differentiable mixing console of neural audio effects
Christian J. Steinmetz, Jordi Pons, Santiago Pascual +1
Applications of deep learning to automatic multitrack mixing are largely unexplored. This is partly due to the limited available data, coupled with the fact that such data is relat…
Less is more: Faster and better music version identification with embedding distillation
Furkan Yesiler, Joan Serrà, Emilia Gómez
Version identification systems aim to detect different renditions of the same underlying musical composition (loosely called cover songs). By learning to encode entire recordings i…
Upsampling artifacts in neural audio synthesis
Jordi Pons, Santiago Pascual, Giulio Cengarle +1
A number of recent advances in neural audio synthesis rely on upsampling layers, which can introduce undesired artifacts. In computer vision, upsampling artifacts have been studied…
SESQA: semi-supervised learning for speech quality assessment
Joan Serrà, Jordi Pons, Santiago Pascual
Automatic speech quality assessment is an important, transversal task whose progress is hampered by the scarcity of human annotations, poor generalization to unseen recording condi…