activity
20172023
most citedTowards Realistic Visual Dubbing with Heterogeneous Sources

34 citations · 114 across the 28 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2020

Improving Accent Conversion with Reference Encoder and End-To-End Text-To-Speech

Wenjie Li, Benlai Tang, Xiang Yin +6

Accent conversion (AC) transforms a non-native speaker's accent into a native accent while maintaining the speaker's voice timbre. In this paper, we propose approaches to improving…

cs.CL20192 cited

A unified sequence-to-sequence front-end model for Mandarin text-to-speech synthesis

Junjie Pan, Xiang Yin, Zhiling Zhang +4

In Mandarin text-to-speech (TTS) system, the front-end text processing module significantly influences the intelligibility and naturalness of synthesized speech. Building a typical…

cs.CL2019

A hybrid text normalization system using multi-head self-attention for mandarin

Junhui Zhang, Junjie Pan, Xiang Yin +5

In this paper, we propose a hybrid text normalization system using multi-head self-attention. The system combines the advantages of a rule-based model and a neural model for text p…

cs.CL20173 cited

Frame Stacking and Retaining for Recurrent Neural Network Acoustic Model

Xu Tian, Jun Zhang, Zejun Ma +2

Frame stacking is broadly applied in end-to-end neural network training like connectionist temporal classification (CTC), and it leads to more accurate models and faster decoding.…

cs.CL201723 cited

Deep LSTM for Large Vocabulary Continuous Speech Recognition

Xu Tian, Jun Zhang, Zejun Ma +6

Recurrent neural networks (RNNs), especially long short-term memory (LSTM) RNNs, are effective network for sequential task like speech recognition. Deeper LSTM models perform well…

cs.CL20173 cited

Exponential Moving Average Model in Parallel Speech Recognition Training

Xu Tian, Jun Zhang, Zejun Ma +2

As training data rapid growth, large-scale parallel training with multi-GPUs cluster is widely applied in the neural network model learning currently.We present a new approach that…