9 papers
Voice conversion with limited data and limitless data augmentations
Olga Slizovskaia, Jordi Janer, Pritish Chandna +1
Applying changes to an input speech signal to change the perceived speaker of speech to a target while maintaining the content of the input is a challenging but interesting task kn…
Locate This, Not That: Class-Conditioned Sound Event DOA Estimation
Olga Slizovskaia, Gordon Wichern, Zhong-Qiu Wang +1
Existing systems for sound event localization and detection (SELD) typically operate by estimating a source location for all classes at every time instant. In this paper, we propos…
Solos: A Dataset for Audio-Visual Music Analysis
Juan F. Montesinos, Olga Slizovskaia, Gloria Haro
In this paper, we present a new dataset of music performance videos which can be used for training machine learning methods for multiple tasks such as audio-visual blind source sep…
Conditioned Source Separation for Music Instrument Performances
Olga Slizovskaia, Gloria Haro, Emilia Gómez
In music source separation, the number of sources may vary for each piece and some of the sources may belong to the same family of instruments, thus sharing timbral characteristics…
Vocoder-Based Speech Synthesis from Silent Videos
Daniel Michelsanti, Olga Slizovskaia, Gloria Haro +3
Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility…
Input complexity and out-of-distribution detection with likelihood-based generative models
Joan Serrà, David Álvarez, Vicenç Gómez +3
Likelihood-based generative models are a promising resource to detect out-of-distribution (OOD) inputs which could compromise the robustness or reliability of a machine learning sy…