2 citations · 3 across the 4 of their papers we have counts for
4 papers
GASS: Generalizing Audio Source Separation with Large-scale Data
Jordi Pons, Xiaoyu Liu, Santiago Pascual +1
Universal source separation targets at separating the audio sources of an arbitrary mix, removing the constraint to operate on a specific domain like speech or music. Yet, the pote…
CLIPSonic: Text-to-Audio Synthesis with Unlabeled Videos and Pretrained Language-Vision Models
Hao-Wen Dong, Xiaoyu Liu, Jordi Pons +5
Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acqu…
Towards Robust Image-in-Audio Deep Steganography
Jaume Ros, Margarita Geleta, Jordi Pons +1
The field of steganography has experienced a surge of interest due to the recent advancements in AI-powered techniques, particularly in the context of multimodal setups that enable…
PodcastMix: A dataset for separating music and speech in podcasts
Nicolás Schmidt, Jordi Pons, Marius Miron
We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. We aim at defining a benchmark suitable for training and e…