activity
20202022
most citedLibri-adhoc40: A dataset collected from synchronized ad-hoc microphone arrays

8 citations · 8 across the 4 of their papers we have counts for

collaborators

7 papers

eess.AS2022

WeKws: A production first small-footprint end-to-end Keyword Spotting Toolkit

Jie Wang, Menglong Xu, Jingyong Hou +4

Keyword spotting (KWS) enables speech-based user interaction and gradually becomes an indispensable component of smart devices. Recently, end-to-end (E2E) methods have become the m…

cs.SD2021

Conformer-based End-to-end Speech Recognition With Rotary Position Embedding

Shengqiang Li, Menglong Xu, Xiao-Lei Zhang

Transformer-based end-to-end speech recognition models have received considerable attention in recent years due to their high training speed and ability to model a long-range globa…

eess.AS2021

AUC Optimization for Robust Small-footprint Keyword Spotting with Limited Training Data

Menglong Xu, Shengqiang Li, Chengdong Liang +1

Deep neural networks provide effective solutions to small-footprint keyword spotting (KWS). However, if training data is limited, it remains challenging to achieve robust and highl…

eess.AS20218 cited

Libri-adhoc40: A dataset collected from synchronized ad-hoc microphone arrays

Shanzheng Guan, Shupei Liu, Junqi Chen +8

Recently, there is a research trend on ad-hoc microphone arrays. However, most research was conducted on simulated data. Although some data sets were collected with a small number…

cs.SD2021

Efficient conformer-based speech recognition with linear attention

Shengqiang Li, Menglong Xu, Xiao-Lei Zhang

Recently, conformer-based end-to-end automatic speech recognition, which outperforms recurrent neural network based ones, has received much attention. Although the parallel computi…

cs.SD2021

Transformer-based end-to-end speech recognition with residual Gaussian-based self-attention

Chengdong Liang, Menglong Xu, Xiao-Lei Zhang

Self-attention (SA), which encodes vector sequences according to their pairwise similarity, is widely used in speech recognition due to its strong context modeling ability. However…