activity
20242026
most citedWenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

1 citations · 1 across the 15 of their papers we have counts for

collaborators

15 papers

cs.CL2026

Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +2

Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing ev…

eess.AS2026

Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation

Jingbin Hu, Qirui Zhan, Yuang Cao +7

We propose a preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressiv…

cs.CL2026

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

Ziyu Zhang, Satoshi Nakamura

Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is de…

cs.SD2026

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang +10

While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the l…

eess.AS2026

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models

Zhixian Zhao, Shuiyuan Wang, Wenjie Tian +3

While Audio Language Models (ALMs) demonstrate strong semantic understanding, they struggle with complex affective interactions. Specifically, textual semantic dominance often over…

eess.AS2026

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions

Chunyu Qiang, Xiaopeng Wang, Kang Yin +11

Generative audio modeling has largely been fragmented into specialized tasks, text-to-speech (TTS), text-to-music (TTM), and text-to-audio (TTA), each operating under heterogeneous…