34 citations · 114 across the 28 of their papers we have counts for
6 papers · 1 filter
Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech
Wenjie Li, Benlai Tang, Xiang Yin +6
Accent conversion (AC) transforms a non-native speaker's accent into a native accent while maintaining the speaker's voice timbre. In this paper, we propose approaches to improving…
A unified sequence-to-sequence front-end model for Mandarin text-to-speech synthesis
Junjie Pan, Xiang Yin, Zhiling Zhang +4
In Mandarin text-to-speech (TTS) system, the front-end text processing module significantly influences the intelligibility and naturalness of synthesized speech. Building a typical…
A hybrid text normalization system using multi-head self-attention for mandarin
Junhui Zhang, Junjie Pan, Xiang Yin +5
In this paper, we propose a hybrid text normalization system using multi-head self-attention. The system combines the advantages of a rule-based model and a neural model for text p…
Frame Stacking and Retaining for Recurrent Neural Network Acoustic Model
Xu Tian, Jun Zhang, Zejun Ma +2
Frame stacking is broadly applied in end-to-end neural network training like connectionist temporal classification (CTC), and it leads to more accurate models and faster decoding.…
Deep LSTM for Large Vocabulary Continuous Speech Recognition
Xu Tian, Jun Zhang, Zejun Ma +6
Recurrent neural networks (RNNs), especially long short-term memory (LSTM) RNNs, are effective network for sequential task like speech recognition. Deeper LSTM models perform well…
Exponential Moving Average Model in Parallel Speech Recognition Training
Xu Tian, Jun Zhang, Zejun Ma +2
As training data rapid growth, large-scale parallel training with multi-GPUs cluster is widely applied in the neural network model learning currently.We present a new approach that…