2 papers
eess.AS2024
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
Xiaoxue Gao, Nancy F. Chen
Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models…
eess.AS2024
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
Xiaoxue Gao, Chen Zhang, Yiming Chen +2
Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a…