2 citations · 7 across the 6 of their papers we have counts for
6 papers
Text-based Talking Video Editing with Cascaded Conditional Diffusion
Bo Han, Heqing Zou, Haoyang Li +2
Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging…
Cross-Modality and Within-Modality Regularization for Audio-Visual DeepFake Detection
Heqing Zou, Meng Shen, Yuchen Hu +3
Audio-visual deepfake detection scrutinizes manipulations in public video using complementary multimodal cues. Current methods, which train on fused multimodal data for multimodal…
UniS-MMC: Multimodal Classification via Unimodality-supervised Multimodal Contrastive Learning
Heqing Zou, Meng Shen, Chen Chen +3
Multimodal learning aims to imitate human beings to acquire complementary information from multiple modalities for various downstream tasks. However, traditional aggregation-based…
Cross-Modal Global Interaction and Local Alignment for Audio-Visual Speech Recognition
Yuchen Hu, Ruizhe Li, Chen Chen +3
Audio-visual speech recognition (AVSR) research has gained a great success recently by improving the noise-robustness of audio-only automatic speech recognition (ASR) with noise-in…
Unsupervised Noise adaptation using Data Simulation
Chen Chen, Yuchen Hu, Heqing Zou +2
Deep neural network based speech enhancement approaches aim to learn a noisy-to-clean transformation using a supervised learning paradigm. However, such a trained-well transformati…
Unifying Speech Enhancement and Separation with Gradient Modulation for End-to-End Noise-Robust Speech Separation
Yuchen Hu, Chen Chen, Heqing Zou +2
Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they woul…