activity
20242026
most citedSpark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

1 citations · 2 across the 16 of their papers we have counts for

collaborators

17 papers

cs.CL2026

Complex-Text Robustness Evaluation and Failure Diagnosis for Low-Resource Multilingual Text-to-Speech

Tianlun Zuo, Ziyu Zhang, Tingzhi Mao +2

Low-resource multilingual text-to-speech (TTS) systems have expanded language coverage, but their robustness under complex text inputs remains insufficiently diagnosed. Existing ev…

eess.AS2026

Preference Optimization with LALM Feedback for Continuous Autoregressive Non-Verbal Vocalization Generation

Jingbin Hu, Qirui Zhan, Yuang Cao +7

We propose a preference optimization framework with Large Audio-Language Model (LALM) feedback for controllable non-verbal vocalization (NVV) generation in continuous autoregressiv…

cs.CL2026

Rubric-Aligned Disentangled Evaluation of Human Simultaneous Interpreting

Ziyu Zhang, Satoshi Nakamura

Human simultaneous interpreting (SI) is commonly assessed with analytic rubrics separating meaning transfer, delivery quality, and temporal synchrony, yet no automatic metric is de…

eess.AS2026

Source-Adaptive Data Curation for Bilingual NVV-Aware ASR

Yuang Cao, Qirui Zhan, Jingbin Hu +7

Nonverbal vocalizations (NVVs), such as laughter, sighs, breaths, and coughs, convey affective and interactional information that conventional automatic speech recognition (ASR) sy…

eess.AS2026

SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling

Haoyu Zhang, Jingbin Hu, Hanke Xie +8

With the rapid development of speech generation technology, discrete codec representations have been widely used because they provide a stable prediction paradigm. In expressive sp…

cs.SD2026

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang +10

While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the l…