1 citations · 1 across the 15 of their papers we have counts for
15 papers
Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech
Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +2
Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing ev…
Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation
Jingbin Hu, Qirui Zhan, Yuang Cao +7
We propose a preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressiv…
Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting
Ziyu Zhang, Satoshi Nakamura
Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is de…
Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation
Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang +10
While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the l…
Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models
Zhixian Zhao, Shuiyuan Wang, Wenjie Tian +3
While Audio Language Models (ALMs) demonstrate strong semantic understanding, they struggle with complex affective interactions. Specifically, textual semantic dominance often over…
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
Chunyu Qiang, Xiaopeng Wang, Kang Yin +11
Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous…