6 papers
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
Xiaoxue Gao, Zexin Li, Yiming Chen +1
The emergence of large-scale automatic speech recognition (ASR) models such as Whisper has greatly expanded their adoption across diverse real-world applications. Ensuring robustne…
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
Xiaoxue Gao, Huayun Zhang, Nancy F. Chen
Generative speech models have demonstrated significant potential in improving human-machine interactions, offering valuable real-world applications such as language learning for ch…
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
Xiaoxue Gao, Huayun Zhang, Nancy F. Chen
Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, mak…
SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning
Zhengyuan Liu, Geyu Lin, Hui Li Tan +8
The integration of generative artificial intelligence into educational applications has enhanced personalized and interactive learning experiences, and it shows strong potential to…
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
Xiaoxue Gao, Nancy F. Chen
Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models…
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
Xiaoxue Gao, Chen Zhang, Yiming Chen +2
Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a…