activity
20202023
most citedPAN: Towards Fast Action Recognition via Learning Persistence of Appearance

32 citations · 110 across the 26 of their papers we have counts for

collaborators

25 papers

cs.CL2023

ML-LMCL: Mutual Learning and Large-Margin Contrastive Learning for Improving ASR Robustness in Spoken Language Understanding

Xuxin Cheng, Bowen Cao, Qichen Ye +3

Spoken language understanding (SLU) is a fundamental task in the task-oriented dialogue systems. However, the inevitable errors from automatic speech recognition (ASR) usually impa…

cs.SD2023

NADiffuSE: Noise-aware Diffusion-based Model for Speech Enhancement

Wen Wang, Dongchao Yang, Qichen Ye +2

The goal of speech enhancement (SE) is to eliminate the background interference from the noisy speech signal. Generative models such as diffusion models (DM) have been applied to t…

cs.CL202214 cited

M3ST: Mix at Three Levels for Speech Translation

Xuxin Cheng, Qianqian Dong, Fengpeng Yue +3

How to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It's well known that data augmentation is an efficient method to improve performance for many…

cs.CV20221 cited

Aligning Source Visual and Target Language Domains for Unpaired Video Captioning

Fenglin Liu, Xian Wu, Chenyu You +3

Training supervised video captioning model requires coupled video-caption pairs. However, for many targeted languages, sufficient paired data are not available. To this end, we int…

cs.CL20226 cited

A Dynamic Graph Interactive Framework with Label-Semantic Injection for Spoken Language Understanding

Zhihong Zhu, Weiyuan Xu, Xuxin Cheng +2

Multi-intent detection and slot filling joint models are gaining increasing traction since they are closer to complicated real-world scenarios. However, existing approaches (1) foc…

cs.SD20221 cited

NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS

Dongchao Yang, Songxiang Liu, Jianwei Yu +3

Expressive text-to-speech (TTS) can synthesize a new speaking style by imiating prosody and timbre from a reference audio, which faces the following challenges: (1) The highly dyna…