activity
20202025
most citedLightweight End-to-end Text-to-speech Synthesis for low resource on-device applications

6 citations · 14 across the 8 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2025

Investigating self-supervised features for expressive, multilingual voice conversion

Álvaro Martín-Cortinas, Daniel Sáez-Trigueros, Grzegorz Beringer +7

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between d…

eess.AS2024

Del Visual al Auditivo: Sonorización de Escenas Guiada por Imagen

María Sánchez, Laura Fernández, Julián Arias +6

Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality w…

eess.AS20222 cited

Cross-speaker style transfer for text-to-speech using data augmentation

Manuel Sam Ribeiro, Julian Roth, Giulia Comini +3

We address the problem of cross-speaker style transfer for text-to-speech (TTS) using data augmentation via voice conversion. We assume to have a corpus of neutral non-expressive d…

eess.AS2021

Enhancing audio quality for expressive Neural Text-to-Speech

Abdelhamid Ezzerg, Adam Gabrys, Bartosz Putrycz +7

Artificial speech synthesis has made a great leap in terms of naturalness as recent Text-to-Speech (TTS) systems are capable of producing speech with similar quality to human recor…

eess.AS20212 cited

Universal Neural Vocoding with Parallel WaveNet

Yunlong Jiao, Adam Gabrys, Georgi Tinchev +3

We present a universal neural vocoder based on Parallel WaveNet, with an additional conditioning network called Audio Encoder. Our universal vocoder offers real-time high-quality s…

eess.AS20204 cited

Parallel WaveNet conditioned on VAE latent vectors

Jonas Rohnke, Tom Merritt, Jaime Lorenzo-Trueba +4

Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically m…