3 papers
cs.SD2023
VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting
Ao Zhang, He Wang, Pengcheng Guo +5
The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and no…
eess.AS2023
The NPU-ASLP System for Audio-Visual Speech Recognition in MISP 2022 Challenge
Pengcheng Guo, He Wang, Bingshen Mu +2
This paper describes our NPU-ASLP system for the Audio-Visual Diarization and Recognition (AVDR) task in the Multi-modal Information based Speech Processing (MISP) 2022 Challenge.…
cs.SD2022
Minimizing Sequential Confusion Error in Speech Command Recognition
Zhanheng Yang, Hang Lv, Xiong Wang +2
Speech command recognition (SCR) has been commonly used on resource constrained devices to achieve hands-free user experience. However, in real applications, confusion among comman…