activity
20182025
most citedParallel WaveNet conditioned on VAE latent vectors

4 citations · 8 across the 8 of their papers we have counts for

collaborators
Showing eess.ASShow all

10 papers · 1 filter

eess.AS2025

Investigating self-supervised features for expressive, multilingual voice conversion

Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Grzegorz Beringer +7

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between d…

eess.AS2023

Comparing normalizing flows and diffusion models for prosody and acoustic modelling in text-to-speech

Guangyan Zhang, Thomas Merritt, Manuel Sam Ribeiro +10

Neural text-to-speech systems are often optimized on L1/L2 losses, which make strong assumptions about the distributions of the target data space. Aiming to improve those assumptio…

eess.AS2022

Remap, warp and attend: Non-parallel many-to-many accent conversion with Normalizing Flows

Abdelhamid Ezzerg, Thomas Merritt, Kayoko Yanagisawa +6

Regional accents of the same language affect not only how words are pronounced (i.e., phonetic content), but also impact prosodic aspects of speech such as speaking rate and intona…

eess.AS20222 cited

Text-free non-parallel many-to-many voice conversion using normalising flows

Thomas Merritt, Abdelhamid Ezzerg, Piotr Biliński +4

Non-parallel voice conversion (VC) is typically achieved using lossy representations of the source speech. However, ensuring only speaker identity information is dropped whilst all…

eess.AS20204 cited

Parallel WaveNet conditioned on VAE latent vectors

Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba +4

Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically m…

eess.AS2020

Low-resource expressive text-to-speech using data augmentation

Goeric Huybrechts, Thomas Merritt, Giulia Comini +3

While recent neural text-to-speech (TTS) systems perform remarkably well, they typically require a substantial amount of recordings from the target speaker reading in the desired s…