activity
20192025
most citedZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMs

16 citations · 58 across the 56 of their papers we have counts for

collaborators

62 papers

cs.SD2025

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model

Xueyuan Chen, Dongchao Yang, Wenxuan Wu +5

Dysarthric speech reconstruction (DSR) aims to convert dysarthric speech into comprehensible speech while maintaining the speaker's identity. Despite significant advancements, exis…

cs.SD2025

UniSep: Universal Target Audio Separation with Language Models at Scale

Yuanyuan Wang, Hangting Chen, Dongchao Yang +7

We propose Universal target audio Separation (UniSep), addressing the separation task on arbitrary mixtures of different types of audio. Distinguished from previous studies, UniSep…

eess.AS2025

Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT

Dongyang Dai, Zhiyong Wu, Shiyin Kang +5

Grapheme-to-phoneme (G2P) conversion serves as an essential component in Chinese Mandarin text-to-speech (TTS) system, where polyphone disambiguation is the core issue. In this pap…

eess.AS2024

AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions

Yuanyuan Wang, Hangting Chen, Dongchao Yang +2

Current Text-to-audio (TTA) models mainly use coarse text descriptions as inputs to generate audio, which hinders models from generating audio with fine-grained control of content…

cs.SD2024

Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models

Weiqin Li, Peiji Yang, Yicheng Zhong +5

Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabili…

cs.SD2024★ 1 cited

CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction

Xueyuan Chen, Dongchao Yang, Dingdong Wang +3

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech. It still suffers from low speaker similarity and poor prosody naturalness. In this pa…