31 citations · 46 across the 4 of their papers we have counts for
4 papers
An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection
Ganglai Wang, Peng Zhang, Lei Xie +3
DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake vi…
Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild
Ganglai Wang, Peng Zhang, Lei Xie +2
Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standin…
Audio-visual speech separation based on joint feature representation with cross-modal attention
Junwen Xiong, Peng Zhang, Lei Xie +3
Multi-modal based speech separation has exhibited a specific advantage on isolating the target character in multi-talker noisy environments. Unfortunately, most of current separati…
Unsupervised Cross-Modal Distillation for Thermal Infrared Tracking
Jingxian Sun, Lichao Zhang, Yufei Zha +4
The target representation learned by convolutional neural networks plays an important role in Thermal Infrared (TIR) tracking. Currently, most of the top-performing TIR trackers ar…