9 citations · 9 across the 2 of their papers we have counts for
8 papers · 1 filter
ADEPT: A Dataset for Evaluating Prosody Transfer
Alexandra Torresquintero, Tian Huey Teh, Christopher G. R. Wallis +6
Text-to-speech is now able to achieve near-human naturalness and research focus has shifted to increasing expressivity. One popular method is to transfer the prosody from a referen…
Ctrl-P: Temporal Control of Prosodic Variation for Speech Synthesis
Devang S Ram Mohan, Vivian Hu, Tian Huey Teh +6
Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to r…
An Overview of Voice Conversion and its Challenges: From Statistical Modeling to Deep Learning
Berrak Sisman, Junichi Yamagishi, Simon King +1
Speaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic cont…
Perception of prosodic variation for speech synthesis using an unsupervised discrete representation of F0
Zack Hodari, Catherine Lai, Simon King
In English, prosody adds a broad range of information to segment sequences, from information structure (e.g. contrast) to stylistic variation (e.g. expression of emotion). However,…
Using generative modelling to produce varied intonation for speech synthesis
Zack Hodari, Oliver Watts, Simon King
Unlike human speakers, typical text-to-speech (TTS) systems are unable to produce multiple distinct renditions of a given sentence. This has previously been addressed by adding exp…
Attentive Filtering Networks for Audio Replay Attack Detection
Cheng-I Lai, Alberto Abad, Korin Richmond +3
An attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the…