5 papers
Multiple Confidence Gates For Joint Training Of SE And ASR
Tianrui Wang, Weibin Zhu, Yingying Gao +2
Joint training of speech enhancement model (SE) and speech recognition model (ASR) is a common solution for robust ASR in noisy environments. SE focuses on improving the auditory q…
Harmonic gated compensation network plus for ICASSP 2022 DNS CHALLENGE
Tianrui Wang, Weibin Zhu, Yingying Gao +3
The harmonic structure of speech is resistant to noise, but the harmonics may still be partially masked by noise. Therefore, we previously proposed a harmonic gated compensation ne…
Detecting Escalation Level from Speech with Transfer Learning and Acoustic-Lexical Information Fusion
Ziang Zhou, Yanze Xu, Ming Li
Textual escalation detection has been widely applied to e-commerce companies' customer service systems to pre-alert and prevent potential conflicts. Similarly, in public areas such…
Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
Fan Yu, Haoneng Luo, Pengcheng Guo +6
Continuous integrate-and-fire (CIF) based models, which use a soft and monotonic alignment mechanism, have been well applied in non-autoregressive (NAR) speech recognition with com…
Identity-Enhanced Network for Facial Expression Recognition
Yanwei Li, Xingang Wang, Shilei Zhang +4
Facial expression recognition is a challenging task, arguably because of large intra-class variations and high inter-class similarities. The core drawback of the existing approache…