39 citations · 86 across the 19 of their papers we have counts for
4 papers · 1 filter
Dynamic Prosody Generation for Speech Synthesis using Linguistics-Driven Acoustic Embedding Selection
Shubhi Tyagi, Marco Nicolis, Jonas Rohnke +2
Recent advances in Text-to-Speech (TTS) have improved quality and naturalness to near-human capabilities when considering isolated sentences. But something which is still lacking i…
Transformation of low-quality device-recorded speech to high-quality speech using improved SEGAN model
Seyyed Saeed Sarfjoo, Xin Wang, Gustav Eje Henter +3
Nowadays vast amounts of speech data are recorded from low-quality recorder devices such as smartphones, tablets, laptops, and medium-quality microphones. The objective of this res…
Using VAEs and Normalizing Flows for One-shot Text-To-Speech Synthesis of Expressive Speech
Vatsal Aggarwal, Marius Cotescu, Nishant Prateek +2
We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement…
In Other News: A Bi-style Text-to-speech Model for Synthesizing Newscaster Voice with Limited Data
Nishant Prateek, Mateusz Łajszczak, Roberto Barra-Chicote +5
Neural text-to-speech synthesis (NTTS) models have shown significant progress in generating high-quality speech, however they require a large quantity of training data. This makes…