4 papers
Audiobook-CC: Controllable Long-context Speech Generation for Multicast Audiobook
Min Liu, JingJing Yin, Xiang Zhang +6
Existing text-to-speech systems predominantly focus on single-sentence synthesis and lack adequate contextual modeling as well as fine-grained performance control capabilities for…
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
Sijing Chen, Yuan Feng, Laipeng He +22
With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin Audi…
GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition
Yu Pan, Yuguang Yang, Heng Lu +2
The continuous evolution of pre-trained speech models has greatly advanced Speech Emotion Recognition (SER). However, current research typically relies on utterance-level emotion l…
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
Yu Pan, Yuguang Yang, Yuheng Huang +6
Despite notable progress, speech emotion recognition (SER) remains challenging due to the intricate and ambiguous nature of speech emotion, particularly in wild world. While curren…