3 papers
cs.SD2022
Fast-U2++: Fast and Accurate End-to-End Speech Recognition in Joint CTC/Attention Frames
Chengdong Liang, Xiao-Lei Zhang, BinBin Zhang +5
Recently, the unified streaming and non-streaming two-pass (U2/U2++) end-to-end model for speech recognition has shown great performance in terms of streaming capability, accuracy…
cs.SD2022
FusionFormer: Fusing Operations in Transformer for Efficient Streaming Speech Recognition
Xingchen Song, Di Wu, Binbin Zhang +8
The recently proposed Conformer architecture which combines convolution with attention to capture both local and global dependencies has become the \textit{de facto} backbone model…
eess.AS2022
WeKws: A production first small-footprint end-to-end Keyword Spotting Toolkit
Jie Wang, Menglong Xu, Jingyong Hou +4
Keyword spotting (KWS) enables speech-based user interaction and gradually becomes an indispensable component of smart devices. Recently, end-to-end (E2E) methods have become the m…