5 papers
Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
Chengdong Liang, Xiao-Lei Zhang, BinBin Zhang +5
Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy…
FusionFormer: Fusing Operations in Transformer for Efficient Streaming Speech Recognition
Xingchen Song, Di Wu, Binbin Zhang +8
The recently proposed Conformer architecture which combines convolution with attention to capture both local and global dependencies has become the \textit{de facto} backbone model…
Non-Autoregressive Transformer ASR with CTC-Enhanced Decoder Input
Xingchen Song, Zhiyong Wu, Yiheng Huang +3
Non-autoregressive (NAR) transformer models have achieved significantly inference speedup but at the cost of inferior accuracy compared to autoregressive (AR) models in automatic s…
Speech-XLNet: Unsupervised Acoustic Model Pretraining For Self-Attention Networks
Xingchen Song, Guangsen Wang, Zhiyong Wu +4
Self-attention network (SAN) can benefit significantly from the bi-directional representation learning through unsupervised pretraining paradigms such as BERT and XLNet. In this pa…
A Random Gossip BMUF Process for Neural Language Modeling
Yiheng Huang, Jinchuan Tian, Lei Han +4
Neural network language model (NNLM) is an essential component of industrial ASR systems. One important challenge of training an NNLM is to leverage between scaling the learning pr…