39 citations · 86 across the 19 of their papers we have counts for
7 papers · 1 filter
Enhancing audio quality for expressive Neural Text-to-Speech
Abdelhamid Ezzerg, Adam Gabrys, Bartosz Putrycz +7
Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recor…
Voicy: Zero-Shot Non-Parallel Voice Conversion in Noisy Reverberant Environments
Alejandro Mottini, Jaime Lorenzo-Trueba, Sri Vishnu Kumar Karlapati +1
Voice Conversion (VC) is a technique that aims to transform the non-linguistic information of a source utterance to change the perceived identity of the speaker. While there is a r…
A learned conditional prior for the VAE acoustic space of a TTS system
Penny Karanasou, Sri Karlapati, Alexis Moinet +5
Many factors influence speech yielding different renditions of a given sentence. Generative models, such as variational autoencoders (VAEs), capture this variability and allow mult…
Weakly-supervised word-level pronunciation error detection in non-native English speech
Daniel Korzekwa, Jaime Lorenzo-Trueba, Thomas Drugman +2
We propose a weakly-supervised model for word-level mispronunciation detection in non-native (L2) English speech. To train this model, phonetically transcribed L2 speech is not req…
Proteno: Text Normalization with Limited Data for Fast Deployment in Text to Speech Systems
Shubhi Tyagi, Antonio Bonafonte, Jaime Lorenzo-Trueba +1
Developing Text Normalization (TN) systems for Text-to-Speech (TTS) on new languages is hard. We propose a novel architecture to facilitate it for multiple languages while using da…
Mispronunciation Detection in Non-native (L2) English with Uncertainty Modeling
Daniel Korzekwa, Jaime Lorenzo-Trueba, Szymon Zaporowski +3
A common approach to the automatic detection of mispronunciation in language learning is to recognize the phonemes produced by a student and compare it to the expected pronunciatio…