4 citations · 10 across the 10 of their papers we have counts for
16 papers
Remap, warp and attend: Non-parallel many-to-many accent conversion with Normalizing Flows
Abdelhamid Ezzerg, Thomas Merritt, Kayoko Yanagisawa +6
Regional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intona…
Stutter-TTS: Controlled Synthesis and Improved Recognition of Stuttered Speech
Xin Zhang, Iván Vallés-Pérez, Andreas Stolcke +5
Stuttering is a speech disorder where the natural flow of speech is interrupted by blocks, repetitions or prolongations of syllables, words and phrases. The majority of existing au…
Prosodic Alignment for off-screen automatic dubbing
Yogesh Virkar, Marcello Federico, Robert Enyedi +1
The goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence. This entails isochrony, i.e., translating the original speech by als…
Text-free non-parallel many-to-many voice conversion using normalising flows
Thomas Merritt, Abdelhamid Ezzerg, Piotr Biliński +4
Non-parallel voice conversion (VC) is typically achieved using lossy representations of the source speech. However, ensuring only speaker identity information is dropped whilst all…
Voice Filter: Few-shot text-to-speech speaker adaptation using voice conversion as a post-processing module
Adam Gabryś, Goeric Huybrechts, Manuel Sam Ribeiro +6
State-of-the-art text-to-speech (TTS) systems require several hours of recorded speech data to generate high-quality synthetic speech. When using reduced amounts of training data,…
SynthASR: Unlocking Synthetic Data for Speech Recognition
Amin Fazel, Wei Yang, Yulan Liu +4
End-to-end (E2E) automatic speech recognition (ASR) models have recently demonstrated superior performance over the traditional hybrid ASR models. Training an E2E ASR model require…