13 citations · 46 across the 11 of their papers we have counts for
5 papers · 1 filter
Speech BERT Embedding For Improving Prosody in Neural TTS
Liping Chen, Yan Deng, Xi Wang +2
This paper presents a speech BERT model to extract embedded prosody information in speech segments for improving the prosody of synthesized speech in neural text-to-speech (TTS). A…
On Addressing Practical Challenges for RNN-Transducer
Rui Zhao, Jian Xue, Jinyu Li +3
In this paper, several works are proposed to address practical challenges for deploying RNN Transducer (RNN-T) based speech recognition system. These challenges are adapting a well…
s-Transformer: Segment-Transformer for Robust Neural Speech Synthesis
Xi Wang, Huaiping Ming, Lei He +1
Neural end-to-end text-to-speech (TTS) , which adopts either a recurrent model, e.g. Tacotron, or an attention one, e.g. Transformer, to characterize a speech utterance, has achiev…
Developing RNN-T Models Surpassing High-Performance Hybrid Models with Customization Capability
Jinyu Li, Rui Zhao, Zhong Meng +8
Because of its streaming nature, recurrent neural network transducer (RNN-T) is a very promising end-to-end (E2E) model that may replace the popular hybrid model for automatic spee…
Modeling Multi-speaker Latent Space to Improve Neural TTS: Quick Enrolling New Speaker and Enhancing Premium Voice
Yan Deng, Lei He, Frank Soong
Neural TTS has shown it can generate high quality synthesized speech. In this paper, we investigate the multi-speaker latent space to improve neural TTS for adapting the system to…