1 citations · 1 across the 2 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2024
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
Jinlong Xue, Yayue Deng, Yicheng Han +2
Recent advances in large language models (LLMs) and development of audio codecs greatly propel the zero-shot TTS. They can synthesize personalized speech with only a 3-second speec…
cs.SD2023★ 1 cited
Frame-level emotional state alignment method for speech emotion recognition
Qifei Li, Yingming Gao, Cong Wang +4
Speech emotion recognition (SER) systems aim to recognize human emotional state during human-computer interaction. Most existing SER systems are trained based on utterance-level la…