activity
20232025
most citedMSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition

3 citations · 6 across the 9 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2025

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

Yu Pan, Yanni Hu, Yuguang Yang +5

Despite great advances, achieving high-fidelity emotional voice conversion (EVC) with flexible and interpretable control remains challenging. This paper introduces ClapFM-EVC, a no…

cs.SD2024

Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Sijing Chen, Yuan Feng, Laipeng He +22

With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin Audi…

cs.SD2024

GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition

Yu Pan, Yuguang Yang, Heng Lu +2

The continuous evolution of pre-trained speech models has greatly advanced Speech Emotion Recognition (SER). However, current research typically relies on utterance-level emotion l…

cs.SD2024★ 2 cited

PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders

Yu Pan, Xiang Zhang, Yuguang Yang +6

Neural speech codecs have recently emerged as a focal point in the fields of speech compression and generation. Despite this progress, achieving high-quality speech reconstruction…

cs.SD2023★ 3 cited

MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition

Yu Pan, Yuguang Yang, Yuheng Huang +6

Despite notable progress, speech emotion recognition (SER) remains challenging due to the intricate and ambiguous nature of speech emotion, particularly in wild world. While curren…