6 citations · 14 across the 8 of their papers we have counts for
6 papers · 1 filter
Investigating self-supervised features for expressive, multilingual voice conversion
Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Grzegorz Beringer +7
Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between d…
Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen
María Sánchez, Laura Fernández, Julián Arias +6
Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality w…
Cross-speaker style transfer for text-to-speech using data augmentation
Manuel Sam Ribeiro, Julian Roth, Giulia Comini +3
We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive d…
Enhancing audio quality for expressive Neural Text-to-Speech
Abdelhamid Ezzerg, Adam Gabrys, Bartosz Putrycz +7
Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recor…
Universal Neural Vocoding with Parallel WaveNet
Yunlong Jiao, Adam Gabrys, Georgi Tinchev +3
We present a universal neural vocoder based on Parallel WaveNet, with an additional conditioning network called Audio Encoder. Our universal vocoder offers real-time high-quality s…
Parallel WaveNet conditioned on VAE latent vectors
Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba +4
Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically m…