activity
20242026
collaborators

6 papers

eess.AS2026

MORE: Multi-Objective Adversarial Attacks on Speech Recognition

Xiaoxue Gao, Zexin Li, Yiming Chen +1

The emergence of large-scale automatic speech recognition (ASR) models such as Whisper has greatly expanded their adoption across diverse real-world applications. Ensuring robustne…

eess.AS2025

MultiGen: Child-Friendly Multilingual Speech Generator with LLMs

Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Generative speech models have demonstrated significant potential in improving human-machine interactions, offering valuable real-world applications such as language learning for ch…

eess.AS2025

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, mak…

cs.CL2025

SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning

Zhengyuan Liu, Geyu Lin, Hui Li Tan +8

The integration of generative artificial intelligence into educational applications has enhanced personalized and interactive learning experiences, and it shows strong potential to…

eess.AS2024

Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models

Xiaoxue Gao, Nancy F. Chen

Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models…

eess.AS2024

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Xiaoxue Gao, Chen Zhang, Yiming Chen +2

Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a…