4 papers
QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
Zhiyu Wu, Jingyi Fang, Yufei Tang +3
Voice Timbre Attribute Detection (vTAD) plays a pivotal role in fine-grained timbre modeling for speech generation tasks. However, it remains challenging due to the inherently subj…
Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations
Jinming Chen, Jingyi Fang, Yuanzhong Zheng +2
Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving th…
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
Yuqi Li, Yuanzhong Zheng, Zhongtian Guo +3
This paper presents SpecWav-Attack, an adversarial model for detecting speakers in anonymized speech. It leverages Wav2Vec2 for feature extraction and incorporates spectrogram resi…
SFE-Net: Harnessing Biological Principles of Differential Gene Expression for Improved Feature Selection in Deep Learning Networks
Yuqi Li, Yuanzhong Zheng, Yaoxuan Wang +2
In the realm of DeepFake detection, the challenge of adapting to various synthesis methodologies such as Faceswap, Deepfakes, Face2Face, and NeuralTextures significantly impacts th…