9 citations · 13 across the 4 of their papers we have counts for
10 papers
Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis
Erica Cooper, Xin Wang, Yi Zhao +2
We explore pretraining strategies including choice of base corpus with the aim of choosing the best strategy for zero-shot multi-speaker end-to-end synthesis. We also examine choic…
How Similar or Different Is Rakugo Speech Synthesizer to Professional Performers?
Shuhei Kato, Yusuke Yasuda, Xin Wang +2
We have been working on speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward speech synthesis that authent…
End-to-End Text-to-Speech using Latent Duration based on VQ-VAE
Yusuke Yasuda, Xin Wang, Junichi Yamagishi
Explicit duration modeling is a key to achieving robust and efficient alignment in text-to-speech synthesis (TTS). We propose a new TTS framework using explicit duration modeling t…
Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis
Yusuke Yasuda, Xin Wang, Junichi Yamagishi
Neural sequence-to-sequence text-to-speech synthesis (TTS) can produce high-quality speech directly from text or simple linguistic features such as phonemes. Unlike traditional pip…
Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?
Erica Cooper, Cheng-I Lai, Yusuke Yasuda +1
Previous work on speaker adaptation for end-to-end speech synthesis still falls short in speaker similarity. We investigate an orthogonal approach to the current speaker adaptation…
Modeling of Rakugo Speech and Its Limitations: Toward Speech Synthesis That Entertains Audiences
Shuhei Kato, Yusuke Yasuda, Xin Wang +3
We have been investigating rakugo speech synthesis as a challenging example of speech synthesis that entertains audiences. Rakugo is a traditional Japanese form of verbal entertain…