2 citations · 2 across the 11 of their papers we have counts for
Showing 2024 · eess.ASShow all
3 papers · 2 filters
eess.AS2024
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
Xiaoxue Gao, Nancy F. Chen
Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models…
eess.AS2024
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
Xiaoxue Gao, Chen Zhang, Yiming Chen +2
Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a…
eess.AS2024
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
Xiaoxue Gao, Yiming Chen, Xianghu Yue +2
Text-to-speech (TTS) has been extensively studied for generating high-quality speech with textual inputs, playing a crucial role in various real-time applications. For real-world d…