activity
20202024
most citedExploring Emotion Features and Fusion Strategies for Audio-Video Emotion Recognition

75 citations · 78 across the 5 of their papers we have counts for

collaborators

5 papers

eess.AS2024

Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment

Li Kewei, Zhou Hengshun, Shen Kai +2

User-defined keyword spotting (KWS) enhances the user experience by allowing individuals to customize keywords. However, in open-vocabulary scenarios, most existing methods commonl…

eess.AS20233 cited

Hierarchical Audio-Visual Information Fusion with Multi-label Joint Decoding for MER 2023

Haotian Wang, Yuxuan Xi, Hang Chen +11

In this paper, we propose a novel framework for recognizing both discrete and dimensional emotions. In our framework, deep features extracted from foundation models are used as rob…

eess.AS2023

The USTC-NERCSLIP Systems for the CHiME-7 DASR Challenge

Ruoyu Wang, Maokui He, Jun Du +16

This technical report details our submission system to the CHiME-7 DASR Challenge, which focuses on speaker diarization and speech recognition under complex multi-speaker scenarios…

cs.SD2022

A Study of Designing Compact Audio-Visual Wake Word Spotting System Based on Iterative Fine-Tuning in Neural Network Pruning

Hengshun Zhou, Jun Du, Chao-Han Huck Yang +2

Audio-only-based wake word spotting (WWS) is challenging under noisy conditions due to environmental interference in signal transmission. In this paper, we investigate on designing…

cs.CV202075 cited

Exploring Emotion Features and Fusion Strategies for Audio-Video Emotion Recognition

Hengshun Zhou, Debin Meng, Yuanyuan Zhang +4

The audio-video based emotion recognition aims to classify a given video into basic emotions. In this paper, we describe our approaches in EmotiW 2019, which mainly explores emotio…