activity
20242026
collaborators
Showing 2024Show all

10 papers · 1 filter

eess.AS2024

Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

Xuechen Wang, Shiwan Zhao, Haoqin Sun +3

Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective m…

cs.SD2024

PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge

Shiyao Wang, Jiaming Zhou, Shiwan Zhao +1

For the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge, we introduce the PB-LRDWWS system. This system combines a dysarthric speech content feature extra…

cs.SD2024

CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition

Tian-Hao Zhang, Dinghao Zhou, Guiping Zhong +2

RNN-T models are widely used in ASR, which rely on the RNN-T loss to achieve length alignment between input audio and target sequence. However, the implementation complexity and th…

eess.AS2024

Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge

Hongfei Xue, Rong Gong, Mingchen Shao +10

The StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recogni…

cs.LG2024

Uncertainty-Aware Mean Opinion Score Prediction

Hui Wang, Shiwan Zhao, Jiaming Zhou +4

Mean Opinion Score (MOS) prediction has made significant progress in specific domains. However, the unstable performance of MOS prediction models across diverse samples presents on…

cs.SD2024

Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition

Haoqin Sun, Shiwan Zhao, Xiangyu Kong +4

Recognizing emotions from speech is a daunting task due to the subtlety and ambiguity of expressions. Traditional speech emotion recognition (SER) systems, which typically rely on…