activity
20232026
most citedNaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

20 citations · 41 across the 24 of their papers we have counts for

collaborators
Showing cs.SDShow all

8 papers · 1 filter

cs.SD2026

UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization

Dongchao Yang, Yuanyuan Wang, Dading Chong +3

We study two foundational problems in audio language models: (1) how to design an audio tokenizer that can serve as an intermediate representation for both understanding and genera…

cs.SD2025

DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models

Yuanyuan Wang, Dongchao Yang, Yiwen Shao +5

Extending pre-trained text Large Language Models (LLMs)'s speech understanding or generation abilities by introducing various effective speech tokens has attracted great attention…

cs.SD2025

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model

Xueyuan Chen, Dongchao Yang, Wenxuan Wu +5

Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, exis…

cs.SD2025

ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling

Dongchao Yang, Songxiang Liu, Haohan Guo +9

Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the ap…

cs.SD2025

UniSep: Universal Target Audio Separation with Language Models at Scale

Yuanyuan Wang, Hangting Chen, Dongchao Yang +7

We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep…

cs.SD2024

Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

Haohan Guo, Fenglong Xie, Dongchao Yang +2

The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by ``recency bias", CLM lacks sufficient attentio…