4 citations · 8 across the 5 of their papers we have counts for
11 papers
Remap, warp and attend: Non-parallel many-to-many accent conversion with Normalizing Flows
Abdelhamid Ezzerg, Thomas Merritt, Kayoko Yanagisawa +6
Regional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intona…
Text-free non-parallel many-to-many voice conversion using normalising flows
Thomas Merritt, Abdelhamid Ezzerg, Piotr Biliński +4
Non-parallel voice conversion (VC) is typically achieved using lossy representations of the source speech. However, ensuring only speaker identity information is dropped whilst all…
Non-Autoregressive TTS with Explicit Duration Modelling for Low-Resource Highly Expressive Speech
Raahil Shah, Kamil Pokora, Abdelhamid Ezzerg +5
Whilst recent neural text-to-speech (TTS) approaches produce high-quality speech, they typically require a large amount of recordings from the target speaker. In previous work, a 3…
Parallel WaveNet conditioned on VAE latent vectors
Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba +4
Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically m…
Low-resource expressive text-to-speech using data augmentation
Goeric Huybrechts, Thomas Merritt, Giulia Comini +3
While recent neural text-to-speech (TTS) systems perform remarkably well, they typically require a substantial amount of recordings from the target speaker reading in the desired s…
CAMP: a Two-Stage Approach to Modelling Prosody in Context
Zack Hodari, Alexis Moinet, Sri Karlapati +6
Prosody is an integral part of communication, but remains an open problem in state-of-the-art speech synthesis. There are two major issues faced when modelling prosody: (1) prosody…