activity
20212024
most citedUnsupervised Cross-Modal Distillation for Thermal Infrared Tracking

31 citations · 52 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

7 papers · 1 filter

cs.CV2024

DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction

Junwen Xiong, Peng Zhang, Tao You +3

Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as…

cs.CV20232 cited

UniST: Towards Unifying Saliency Transformer for Video Saliency Prediction and Detection

Junwen Xiong, Peng Zhang, Chuanyue Li +3

Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceiving dynamic s…

cs.CV20234 cited

Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization

Tianyu Liu, Peng Zhang, Wei Huang +3

Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to esta…

cs.CV2023

FTFDNet: Learning to Detect Talking Face Video Manipulation with Tri-Modality Interaction

Ganglai Wang, Peng Zhang, Junwen Xiong +3

DeepFake based digital facial forgery is threatening public media security, especially when lip manipulation has been used in talking face generation, and the difficulty of fake vi…

cs.CV20224 cited

An Audio-Visual Attention Based Multimodal Network for Fake Talking Face Videos Detection

Ganglai Wang, Peng Zhang, Lei Xie +3

DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake vi…

cs.CV20229 cited

Attention-Based Lip Audio-Visual Synthesis for Talking Face Generation in the Wild

Ganglai Wang, Peng Zhang, Lei Xie +2

Talking face generation with great practical significance has attracted more attention in recent audio-visual studies. How to achieve accurate lip synchronization is a long-standin…