75 citations · 78 across the 5 of their papers we have counts for
5 papers
Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment
Li Kewei, Zhou Hengshun, Shen Kai +2
User-defined keyword spotting (KWS) enhances the user experience by allowing individuals to customize keywords. However, in open-vocabulary scenarios, most existing methods commonl…
Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023
Haotian Wang, Yuxuan Xi, Hang Chen +11
In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as rob…
The USTC-NERCSLIP Systems for the CHiME-7 DASR Challenge
Ruoyu Wang, Maokui He, Jun Du +16
This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios…
A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network Pruning
Hengshun Zhou, Jun Du, Chao-Han Huck Yang +2
Audio-only-based wake word spotting (WWS) is challenging under noisy conditions due to environmental interference in signal transmission. In this paper, we investigate on designing…
Exploring Emotion Features and Fusion Strategies for Audio-Video Emotion Recognition
Hengshun Zhou, Debin Meng, Yuanyuan Zhang +4
The audio-video based emotion recognition aims to classify a given video into basic emotions. In this paper, we describe our approaches in EmotiW 2019, which mainly explores emotio…