31 citations · 52 across the 7 of their papers we have counts for
7 papers · 1 filter
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
Junwen Xiong, Peng Zhang, Tao You +3
Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as…
UniST: Towards Unifying Saliency Transformer for Video Saliency Prediction and Detection
Junwen Xiong, Peng Zhang, Chuanyue Li +3
Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceiving dynamic s…
Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization
Tianyu Liu, Peng Zhang, Wei Huang +3
Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to esta…
FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction
Ganglai Wang, Peng Zhang, Junwen Xiong +3
DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake vi…
An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection
Ganglai Wang, Peng Zhang, Lei Xie +3
DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake vi…
Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild
Ganglai Wang, Peng Zhang, Lei Xie +2
Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standin…