2 citations · 3 across the 2 of their papers we have counts for
3 papers
cs.CV2022★ 1 cited
Frame-Subtitle Self-Supervision for Multi-Modal Video Question Answering
Jiong Wang, Zhou Zhao, Weike Jin
Multi-modal video question answering aims to predict correct answer and localize the temporal boundary relevant to the question. The temporal annotations of questions improve QA pe…
cs.CV2022★ 2 cited
VLAD-VSA: Cross-Domain Face Presentation Attack Detection with Vocabulary Separation and Adaptation
Jiong Wang, Zhou Zhao, Weike Jin +5
For face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The repre…
cs.CV2018
Attention-based Pyramid Aggregation Network for Visual Place Recognition
Yingying Zhu, Jiong Wang, Lingxi Xie +1
Visual place recognition is challenging in the urban environment and is usually viewed as a large scale image retrieval task. The intrinsic challenges in place recognition exist th…