activity
20182020
most citedZero-Shot Multi-Speaker Text-To-Speech with State-of-the-art Neural Speaker Embeddings

9 citations · 13 across the 4 of their papers we have counts for

collaborators

10 papers

cs.SD20201 cited

Pretraining Strategies, Waveform Model Choice, and Acoustic Configurations for Multi-Speaker End-to-End Speech Synthesis

Erica Cooper, Xin Wang, Yi Zhao +2

We explore pretraining strategies including choice of base corpus with the aim of choosing the best strategy for zero-shot multi-speaker end-to-end synthesis. We also examine choic…

eess.AS2020

How Similar or Different Is Rakugo Speech Synthesizer to Professional Performers?

Shuhei Kato, Yusuke Yasuda, Xin Wang +2

We have been working on speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward speech synthesis that authent…

eess.AS2020

End-to-End Text-to-Speech using Latent Duration based on VQ-VAE

Yusuke Yasuda, Xin Wang, Junichi Yamagishi

Explicit duration modeling is a key to achieving robust and efficient alignment in text-to-speech synthesis (TTS). We propose a new TTS framework using explicit duration modeling t…

eess.AS2020

Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis

Yusuke Yasuda, Xin Wang, Junichi Yamagishi

Neural sequence-to-sequence text-to-speech synthesis (TTS) can produce high-quality speech directly from text or simple linguistic features such as phonemes. Unlike traditional pip…

eess.AS20203 cited

Can Speaker Augmentation Improve Multi-Speaker End-to-End TTS?

Erica Cooper, Cheng-I Lai, Yusuke Yasuda +1

Previous work on speaker adaptation for end-to-end speech synthesis still falls short in speaker similarity. We investigate an orthogonal approach to the current speaker adaptation…

eess.AS2019

Modeling of Rakugo Speech and Its Limitations: Toward Speech Synthesis That Entertains Audiences

Shuhei Kato, Yusuke Yasuda, Xin Wang +3

We have been investigating rakugo speech synthesis as a challenging example of speech synthesis that entertains audiences. Rakugo is a traditional Japanese form of verbal entertain…