8 citations · 8 across the 4 of their papers we have counts for
6 papers
Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
Chengdong Liang, Xiao-Lei Zhang, BinBin Zhang +5
Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy…
Conformer-based End-to-end Speech Recognition With Rotary Position Embedding
Shengqiang Li, Menglong Xu, Xiao-Lei Zhang
Transformer-based end-to-end speech recognition models have received considerable attention in recent years due to their high training speed and ability to model a long-range globa…
AUC Optimization for Robust Small-footprint Keyword Spotting with Limited Training Data
Menglong Xu, Shengqiang Li, Chengdong Liang +1
Deep neural networks provide effective solutions to small-footprint keyword spotting (KWS). However, if training data is limited, it remains challenging to achieve robust and highl…
Libri-adhoc40: A dataset collected from synchronized ad-hoc microphone arrays
Shanzheng Guan, Shupei Liu, Junqi Chen +8
Recently, there is a research trend on ad-hoc microphone arrays. However, most research was conducted on simulated data. Although some data sets were collected with a small number…
Efficient conformer-based speech recognition with linear attention
Shengqiang Li, Menglong Xu, Xiao-Lei Zhang
Recently, conformer-based end-to-end automatic speech recognition, which outperforms recurrent neural network based ones, has received much attention. Although the parallel computi…
Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention
Menglong Xu, Shengqiang Li, Xiao-Lei Zhang
Recently, several studies reported that dot-product selfattention (SA) may not be indispensable to the state-of-theart Transformer models. Motivated by the fact that dense synthesi…