activity
20182021
most citedUsing previous acoustic context to improve Text-to-Speech synthesis

9 citations · 9 across the 2 of their papers we have counts for

collaborators

10 papers

eess.AS2021

ADEPT: A Dataset for Evaluating Prosody Transfer

Alexandra Torresquintero, Tian Huey Teh, Christopher G. R. Wallis +6

Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a referen…

eess.AS2021

Ctrl-P: Temporal Control of Prosodic Variation for Speech Synthesis

Devang S Ram Mohan, Vivian Hu, Tian Huey Teh +6

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to r…

cs.CL20209 cited

Using previous acoustic context to improve Text-to-Speech synthesis

Pilar Oplustil-Gallegos, Simon King

Many speech synthesis datasets, especially those derived from audiobooks, naturally comprise sequences of utterances. Nevertheless, such data are commonly treated as individual, un…

eess.AS2020

An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning

Berrak Sisman, Junichi Yamagishi, Simon King +1

Speaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic cont…

eess.AS2020

Perception of prosodic variation for speech synthesis using an unsupervised discrete representation of F0

Zack Hodari, Catherine Lai, Simon King

In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However,…

cs.CL2020

Comparison of Speech Representations for Automatic Quality Estimation in Multi-Speaker Text-to-Speech Synthesis

Jennifer Williams, Joanna Rownicka, Pilar Oplustil +1

We aim to characterize how different speakers contribute to the perceived output quality of multi-speaker Text-to-Speech (TTS) synthesis. We automatically rate the quality of TTS u…