16 citations · 101 across the 22 of their papers we have counts for
4 papers · 1 filter
Improving Mandarin End-to-End Speech Recognition with Word N-gram Language Model
Jinchuan Tian, Jianwei Yu, Chao Weng +2
Despite the rapid progress of end-to-end (E2E) automatic speech recognition (ASR), it has been shown that incorporating external language models (LMs) into the decoding can further…
Synthesising Expressiveness in Peking Opera via Duration Informed Attention Network
Yusong Wu, Shengchen Li, Chengzhu Yu +4
This paper presents a method that generates expressive singing voice of Peking opera. The synthesis of expressive opera singing usually requires pitch contours to be extracted as t…
Minimum Bayes Risk Training of RNN-Transducer for End-to-End Speech Recognition
Chao Weng, Chengzhu Yu, Jia Cui +2
In this work, we propose minimum Bayes risk (MBR) training of RNN-Transducer (RNN-T) for end-to-end speech recognition. Specifically, initialized with a RNN-T trained model, MBR tr…
DurIAN: Duration Informed Attention Network For Multimodal Synthesis
Chengzhu Yu, Heng Lu, Na Hu +9
In this paper, we present a generic and robust multimodal synthesis system that produces highly natural speech and facial expression simultaneously. The key component of this syste…