11 citations · 26 across the 10 of their papers we have counts for
8 papers · 1 filter
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Jianwei Yu +3
Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…
Masked Spectrogram Prediction For Self-Supervised Audio Pre-Training
Dading Chong, Helin Wang, Peilin Zhou +1
Transformer-based models attain excellent results and generalize well when trained on sufficient amounts of data. However, constrained by the limited data available in the audio do…
RaDur: A Reference-aware and Duration-robust Network for Target Sound Detection
Dongchao Yang, Helin Wang, Zhongjie Ye +2
Target sound detection (TSD) aims to detect the target sound from a mixture audio given the reference information. Previous methods use a conditional network to extract a sound-dis…
Improving Target Sound Extraction with Timestamp Information
Helin Wang, Dongchao Yang, Chao Weng +2
Target sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events. The previous works mainly focus on the p…
Improving the Performance of Automated Audio Captioning via Integrating the Acoustic and Semantic Information
Zhongjie Ye, Helin Wang, Dongchao Yang +1
Automated audio captioning (AAC) has developed rapidly in recent years, involving acoustic signal processing and natural language processing to generate human-readable sentences fo…
Unsupervised Multi-Target Domain Adaptation for Acoustic Scene Classification
Dongchao Yang, Helin Wang, Yuexian Zou
It is well known that the mismatch between training (source) and test (target) data distribution will significantly decrease the performance of acoustic scene classification (ASC)…