activity
20242026
most citedCosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

4 citations · 4 across the 4 of their papers we have counts for

collaborators

5 papers

eess.AS2026

StepAudio 2.5 Technical Report

Bin Lin, Bo Zhao, Boyong Wu +98

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks.…

cs.SD2025

JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis

Fan Yu, Tao Wang, You Wu +22

Large speech generation models are evolving from single-speaker, short sentence synthesis to multi-speaker, long conversation geneartion. Current long-form speech generation models…

cs.SD2025

SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer

Zhengyan Sheng, Zhihao Du, Shiliang Zhang +2

Current text-to-speech (TTS) models face a persistent limitation: autoregressive (AR) models suffer from low generation efficiency, while modern non-autoregressive (NAR) models exp…

cs.SD2025

Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation

Zhengyan Sheng, Zhihao Du, Heng Lu +2

Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generat…

cs.SD2024★ 4 cited

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Zhihao Du, Yuxuan Wang, Qian Chen +16

In our previous work, we introduced CosyVoice, a multilingual speech synthesis model based on supervised discrete speech tokens. By employing progressive semantic decoding with two…