6 citations · 6 across the 1 of their papers we have counts for
5 papers
Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
Xingchen Song, Zhiyong Wu, Yiheng Huang +3
Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic s…
Masked Pre-trained Encoder base on Joint CTC-Transformer
Lu Liu, Yiheng Huang
This study (The work was accomplished during the internship in Tencent AI lab) addresses semi-supervised acoustic modeling, i.e. attaining high-level representations from unsupervi…
Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks
Xingchen Song, Guangsen Wang, Zhiyong Wu +4
Self-attention network (SAN) can benefit significantly from the bi-directional representation learning through unsupervised pretraining paradigms such as BERT and XLNet. In this pa…
Phrase-Level Class based Language Model for Mandarin Smart Speaker Query Recognition
Yiheng Huang, Liqiang He, Lei Han +2
The success of speech assistants requires precise recognition of a number of entities on particular contexts. A common solution is to train a class-based n-gram language model and…
A Random Gossip BMUF Process for Neural Language Modeling
Yiheng Huang, Jinchuan Tian, Lei Han +4
Neural network language model (NNLM) is an essential component of industrial ASR systems. One important challenge of training an NNLM is to leverage between scaling the learning pr…