activity
20202022
most citedImproving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information

11 citations · 13 across the 7 of their papers we have counts for

collaborators

8 papers

cs.SD20221 cited

NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS

Dongchao Yang, Songxiang Liu, Jianwei Yu +3

Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…

eess.AS2022

Speaker-Aware Mixture of Mixtures Training for Weakly Supervised Speaker Extraction

Zifeng Zhao, Rongzhi Gu, Dongchao Yang +2

Dominant researches adopt supervised training for speaker extraction, while the scarcity of ideally clean corpus and channel mismatch problem are rarely considered. To this end, we…

cs.SD2022

RaDur: A Reference-aware and Duration-robust Network for Target Sound Detection

Dongchao Yang, Helin Wang, Zhongjie Ye +2

Target sound detection (TSD) aims to detect the target sound from a mixture audio given the reference information. Previous methods use a conditional network to extract a sound-dis…

eess.AS2022

Target Confusion in End-to-end Speaker Extraction: Analysis and Approaches

Zifeng Zhao, Dongchao Yang, Rongzhi Gu +2

Recently, end-to-end speaker extraction has attracted increasing attention and shown promising results. However, its performance is often inferior to that of a blind source separat…

cs.SD2022

Improving Target Sound Extraction with Timestamp Information

Helin Wang, Dongchao Yang, Chao Weng +2

Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the p…

cs.SD202111 cited

Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information

Zhongjie Ye, Helin Wang, Dongchao Yang +1

Automated audio captioning (AAC) has developed rapidly in recent years, involving acoustic signal processing and natural language processing to generate human-readable sentences fo…