most citedWenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

1 citations · 1 across the 5 of their papers we have counts for

collaborators

8 papers

cs.SD2026

Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation

Ziyu Zhang, Mingchen Shao, Wenjie Tian +3

Automating clinical documentation from long-form doctor-patient conversations remains challenging for modern audio-language models. While cascaded ASR systems perform well, end-to-…

cs.SD2026

DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis

Ziyu Zhang, Tianlun Zuo, Hanzhao Li +2

Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, they still suffer from limited d…

cs.CL2026

Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +2

Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing ev…

cs.SD2026

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang +10

While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the l…

cs.CL2025

WenetSpeech-Chuan: A Large-Scale Sichuanese Corpus with Rich Annotation for Dialectal Speech Processing

Yuhang Dai, Ziyu Zhang, Shuai Wang +13

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects…

cs.SD2025★ 1 cited

WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation

Longhao Li, Zhao Guo, Hongjie Chen +15

The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS…