activity
20182022
most citedFeature reinforcement with word embedding and parsing information in neural TTS

13 citations · 46 across the 11 of their papers we have counts for

collaborators

16 papers

cs.SD20222 cited

Joint Pre-Training with Speech and Bilingual Text for Direct Speech to Speech Translation

Kun Wei, Long Zhou, Ziqiang Zhang +5

Direct speech-to-speech translation (S2ST) is an attractive research topic with many advantages compared to cascaded S2ST. However, direct S2ST suffers from the data scarcity probl…

cs.SD20226 cited

Cross-Lingual Text-to-Speech Using Multi-Task Learning and Speaker Classifier Joint Training

J. Yang, Lei He

In cross-lingual speech synthesis, the speech in various languages can be synthesized for a monoglot speaker. Normally, only the data of monoglot speakers are available for model t…

cs.SD20212 cited

Cross-speaker Style Transfer with Prosody Bottleneck in Neural Speech Synthesis

Shifeng Pan, Lei He

Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expres…

eess.AS2021

Speech BERT Embedding For Improving Prosody in Neural TTS

Liping Chen, Yan Deng, Xi Wang +2

This paper presents a speech BERT model to extract embedded prosody information in speech segments for improving the prosody of synthesized speech in neural text-to-speech (TTS). A…

eess.AS2021

On Addressing Practical Challenges for RNN-Transducer

Rui Zhao, Jian Xue, Jinyu Li +3

In this paper, several works are proposed to address practical challenges for deploying RNN Transducer (RNN-T) based speech recognition system. These challenges are adapting a well…

cs.CL20214 cited

Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation

Fengpeng Yue, Yan Deng, Lei He +1

Machine Speech Chain, which integrates both end-to-end (E2E) automatic speech recognition (ASR) and text-to-speech (TTS) into one circle for joint training, has been proven to be e…