5 citations · 7 across the 4 of their papers we have counts for
4 papers
Language Adaptive Cross-lingual Speech Representation Learning with Sparse Sharing Sub-networks
Yizhou Lu, Mingkun Huang, Xinghua Qu +2
Unsupervised cross-lingual speech representation learning (XLSR) has recently shown promising results in speech recognition by leveraging vast amounts of unlabeled data across mult…
Improving RNN transducer with normalized jointer network
Mingkun Huang, Jun Zhang, Meng Cai +5
Recurrent neural transducer (RNN-T) is a promising end-to-end (E2E) model in automatic speech recognition (ASR). It has shown superior performance compared to traditional hybrid AS…
Dynamic latency speech recognition with asynchronous revision
Mingkun Huang, Meng Cai, Jun Zhang +4
In this work we propose an inference technique, asynchronous revision, to unify streaming and non-streaming speech recognition models. Specifically, we achieve dynamic latency with…
Modular End-to-end Automatic Speech Recognition Framework for Acoustic-to-word Model
Qi Liu, Zhehuai Chen, Hao Li +3
End-to-end (E2E) systems have played a more and more important role in automatic speech recognition (ASR) and achieved great performance. However, E2E systems recognize output word…