32 citations · 110 across the 26 of their papers we have counts for
25 papers
ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding
Xuxin Cheng, Bowen Cao, Qichen Ye +3
Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impa…
NADiffuSE: Noise-aware Diffusion-based Model for Speech Enhancement
Wen Wang, Dongchao Yang, Qichen Ye +2
The goal of speech enhancement (SE) is to eliminate the background interference from the noisy speech signal. Generative models such as diffusion models (DM) have been applied to t…
M3ST: Mix at Three Levels for Speech Translation
Xuxin Cheng, Qianqian Dong, Fengpeng Yue +3
How to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It's well known that data augmentation is an efficient method to improve performance for many…
Aligning Source Visual and Target Language Domains for Unpaired Video Captioning
Fenglin Liu, Xian Wu, Chenyu You +3
Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we int…
A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding
Zhihong Zhu, Weiyuan Xu, Xuxin Cheng +2
Multi-intent detection and slot filling joint models are gaining increasing traction since they are closer to complicated real-world scenarios. However, existing approaches (1) foc…
NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Jianwei Yu +3
Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…