9 papers
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models
Yizhou Peng, Yukun Ma, Chong Zhang +4
While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion confl…
MMAE: A Massive Multitask Audio Editing Benchmark
Ziyang Ma, Ruiqi Yan, Ruiyang Xu +35
We introduce MMAE, a Massive Multitask Audio Editing benchmark, serving as the first comprehensive evaluation testbed designed for general-purpose instruction-based audio editing.…
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
Yizhou Peng, Ziyang Ma, Changsong Liu +3
Cascaded Automatic Speech Recognition -- Large Language Model (ASR-LLM) pipelines remain popular for industrial Spoken Dialogue Systems (SDS), primarily because their decoupled des…
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
Tianrui Wang, Ziyang Ma, Yizhou Peng +26
Evaluating expressive speech remains challenging, as existing methods mainly assess emotional intensity and overlook whether a speech sample is expressively appropriate for its con…
Evaluating Bias in Spoken Dialogue LLMs for Real-World Decisions and Recommendations
Yihao Wu, Tianrui Wang, Yizhou Peng +5
While biases in large language models (LLMs), such as stereotypes and cultural tendencies in outputs, have been examined and identified, their presence and characteristics in spoke…
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
Yizhou Peng, Yi-Wen Chao, Dianwen Ng +4
Full-duplex spoken dialogue systems (FDSDS) enable more natural human-machine interactions by allowing real-time user interruptions and backchanneling, compared to traditional SDS…