most citedCross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

2 citations · 6 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD20231 cited

Study of GANs for Noisy Speech Simulation from Clean Speech

Leander Melroy Maben, Zixun Guo, Chen Chen +2

The performance of speech processing models trained on clean speech drops significantly in noisy conditions. Training with noisy datasets alleviates the problem, but procuring such…

eess.AS20232 cited

Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition

Yuchen Hu, Ruizhe Li, Chen Chen +3

Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-in…

cs.SD20232 cited

Unsupervised Noise adaptation using Data Simulation

Chen Chen, Yuchen Hu, Heqing Zou +2

Deep neural network based speech enhancement approaches aim to learn a noisy-to-clean transformation using a supervised learning paradigm. However, such a trained-well transformati…

eess.AS20231 cited

Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation

Yuchen Hu, Chen Chen, Heqing Zou +2

Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they woul…

eess.AS2023

Gradient Remedy for Multi-Task Learning in End-to-End Noise-Robust Speech Recognition

Yuchen Hu, Chen Chen, Ruizhe Li +2

Speech enhancement (SE) is proved effective in reducing noise from noisy speech signals for downstream automatic speech recognition (ASR), where multi-task learning strategy is emp…