activity
20202026
most citedMake-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

47 citations · 133 across the 27 of their papers we have counts for

collaborators
Showing 2022 · cs.SDShow all

5 papers · 2 filters

cs.SD2022★ 1 cited

NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS

Dongchao Yang, Songxiang Liu, Jianwei Yu +3

Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…

cs.SD2022★ 9 cited

Diffsound: Discrete Diffusion Model for Text-to-sound Generation

Dongchao Yang, Jianwei Yu, Helin Wang +4

Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound co…

cs.SD2022

RaDur: A Reference-aware and Duration-robust Network for Target Sound Detection

Dongchao Yang, Helin Wang, Zhongjie Ye +2

Target sound detection (TSD) aims to detect the target sound from a mixture audio given the reference information. Previous methods use a conditional network to extract a sound-dis…

cs.SD2022

Improving Target Sound Extraction with Timestamp Information

Helin Wang, Dongchao Yang, Chao Weng +2

Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the p…

cs.SD2022

A Mixed supervised Learning Framework for Target Sound Detection

Dongchao Yang, Helin Wang, Yuexian Zou +1

Target sound detection (TSD) aims to detect the target sound from mixture audio given the reference information. Previous works have shown that TSD models can be trained on fully-a…