activity
20242026
collaborators
Showing 2024Show all

5 papers · 1 filter

eess.AS2024

Transferable Adversarial Attacks against ASR

Xiaoxue Gao, Zexin Li, Yiming Chen +2

Given the extensive research and real-world applications of automatic speech recognition (ASR), ensuring the robustness of ASR models against minor input perturbations becomes a cr…

cs.CL2024

VoiceBench: Benchmarking LLM-Based Voice Assistants

Yiming Chen, Xianghu Yue, Chen Zhang +3

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering…

eess.AS2024

Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization

Xiaoxue Gao, Chen Zhang, Yiming Chen +2

Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a…

cs.SD2024

Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

Yiming Chen, Xianghu Yue, Xiaoxue Gao +4

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primaril…

eess.AS2024

TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations

Xiaoxue Gao, Yiming Chen, Xianghu Yue +2

Text-to-speech (TTS) has been extensively studied for generating high-quality speech with textual inputs, playing a crucial role in various real-time applications. For real-world d…