activity
20232025
most citedMs-senet: Enhancing Speech Emotion Recognition Through Multi-scale Feature Fusion With Squeeze-and-excitation Blocks

3 citations · 4 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2025

QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection

Zhiyu Wu, Jingyi Fang, Yufei Tang +3

Voice Timbre Attribute Detection (vTAD) plays a pivotal role in fine-grained timbre modeling for speech generation tasks. However, it remains challenging due to the inherently subj…

cs.SD2025

SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech

Yuqi Li, Yuanzhong Zheng, Zhongtian Guo +3

This paper presents SpecWav-Attack, an adversarial model for detecting speakers in anonymized speech. It leverages Wav2Vec2 for feature extraction and incorporates spectrogram resi…

eess.AS2025

Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations

Jinming Chen, Jingyi Fang, Yuanzhong Zheng +2

Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving th…

cs.MM2024

SFE-Net: Harnessing Biological Principles of Differential Gene Expression for Improved Feature Selection in Deep Learning Networks

Yuqi Li, Yuanzhong Zheng, Yaoxuan Wang +2

In the realm of DeepFake detection, the challenge of adapting to various synthesis methodologies such as Faceswap, Deepfakes, Face2Face, and NeuralTextures significantly impacts th…

cs.SD2024★ 1 cited

Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition

Jinming Chen, Jingyi Fang, Yuanzhong Zheng +2

Currently, end-to-end (E2E) speech recognition methods have achieved promising performance. However, auto speech recognition (ASR) models still face challenges in recognizing multi…

cs.SD2023★ 3 cited

Ms-senet: Enhancing Speech Emotion Recognition Through Multi-scale Feature Fusion With Squeeze-and-excitation Blocks

Mengbo Li, Yuanzhong Zheng, Dichucheng Li +3

Speech Emotion Recognition (SER) has become a growing focus of research in human-computer interaction. Spatiotemporal features play a crucial role in SER, yet current research lack…