52 citations · 119 across the 3 of their papers we have counts for
5 papers
ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang, Junichi Yamagishi, Massimiliano Todisco +37
Automatic speaker verification (ASV) is one of the most natural and convenient means of biometric person recognition. Unfortunately, just like all other biometric systems, ASV is v…
Evaluating Long-form Text-to-Speech: Comparing the Ratings of Sentences and Paragraphs
Rob Clark, Hanna Silen, Tom Kenter +1
Text-to-speech systems are typically evaluated on single sentences. When long-form content, such as data consisting of full paragraphs or dialogues is considered, evaluating senten…
CHiVE: Varying Prosody in Speech Synthesis with a Linguistically Driven Dynamic Hierarchical Conditional Variational Network
Vincent Wan, Chun-an Chan, Tom Kenter +2
The prosodic aspects of speech signals produced by current text-to-speech systems are typically averaged over training material, and as such lack the variety and liveliness found i…
LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Heiga Zen, Viet Dang, Rob Clark +5
This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use. It is derived from the original audio and text materials of the LibriSpeech corpus, whi…
Uncovering Latent Style Factors for Expressive Speech Synthesis
Yuxuan Wang, RJ Skerry-Ryan, Ying Xiao +5
Prosodic modeling is a core problem in speech synthesis. The key challenge is producing desirable prosody from textual input containing only phonetic information. In this prelimina…