activity
20182022
most citedVoice Conversion for Whispered Speech Synthesis

31 citations · 39 across the 11 of their papers we have counts for

collaborators

25 papers

eess.AS2022

Voice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module

Adam Gabryś, Goeric Huybrechts, Manuel Sam Ribeiro +6

State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data,…

eess.AS20222 cited

Cross-speaker style transfer for text-to-speech using data augmentation

Manuel Sam Ribeiro, Julian Roth, Giulia Comini +3

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive d…

eess.AS2021

Enhancing audio quality for expressive Neural Text-to-Speech

Abdelhamid Ezzerg, Adam Gabrys, Bartosz Putrycz +7

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recor…

cs.SD2021

Voicy: Zero-Shot Non-Parallel Voice Conversion in Noisy Reverberant Environments

Alejandro Mottini, Jaime Lorenzo-Trueba, Sri Vishnu Kumar Karlapati +1

Voice Conversion (VC) is a technique that aims to transform the non-linguistic information of a source utterance to change the perceived identity of the speaker. While there is a r…

eess.AS2021

A learned conditional prior for the VAE acoustic space of a TTS system

Penny Karanasou, Sri Karlapati, Alexis Moinet +5

Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow mult…

eess.AS2021

Weakly-supervised word-level pronunciation error detection in non-native English speech

Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman +2

We propose a weakly-supervised model for word-level mispronunciation detection in non-native (L2) English speech. To train this model, phonetically transcribed L2 speech is not req…