4 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CV2024
DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction
Junwen Xiong, Peng Zhang, Tao You +3
Audio-visual saliency prediction can draw support from diverse modality complements, but further performance enhancement is still challenged by customized architectures as well as…
cs.CV2023★ 2 cited
UniST: Towards Unifying Saliency Transformer for Video Saliency Prediction and Detection
Junwen Xiong, Peng Zhang, Chuanyue Li +3
Video saliency prediction and detection are thriving research domains that enable computers to simulate the distribution of visual attention akin to how humans perceiving dynamic s…
cs.CV2023★ 4 cited
Induction Network: Audio-Visual Modality Gap-Bridging for Self-Supervised Sound Source Localization
Tianyu Liu, Peng Zhang, Wei Huang +3
Self-supervised sound source localization is usually challenged by the modality inconsistency. In recent studies, contrastive learning based strategies have shown promising to esta…