activity
20242026
collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

MORE: Multi-Objective Adversarial Attacks on Speech Recognition

Xiaoxue Gao, Zexin Li, Yiming Chen +1

The emergence of large-scale automatic speech recognition (ASR) models such as Whisper has greatly expanded their adoption across diverse real-world applications. Ensuring robustne…

eess.AS2025

MultiGen: Child-Friendly Multilingual Speech Generator with LLMs

Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Generative speech models have demonstrated significant potential in improving human-machine interactions, offering valuable real-world applications such as language learning for ch…

eess.AS2025

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions

Xiaoxue Gao, Huayun Zhang, Nancy F. Chen

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, mak…

eess.AS2024

Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models

Xiaoxue Gao, Nancy F. Chen

Current automatic speech recognition systems struggle with modeling long speech sequences due to high quadratic complexity of Transformer-based models. Selective state space models…

eess.AS2024

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Xiaoxue Gao, Chen Zhang, Yiming Chen +2

Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a…

eess.AS2024

TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations

Xiaoxue Gao, Yiming Chen, Xianghu Yue +2

Text-to-speech (TTS) has been extensively studied for generating high-quality speech with textual inputs, playing a crucial role in various real-time applications. For real-world d…