2 citations · 3 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 1 cited
Frame-Subtitle Self-Supervision for Multi-Modal Video Question Answering
Jiong Wang, Zhou Zhao, Weike Jin
Multi-modal video question answering aims to predict correct answer and localize the temporal boundary relevant to the question. The temporal annotations of questions improve QA pe…
cs.CV2022★ 2 cited
VLAD-VSA: Cross-Domain Face Presentation Attack Detection with Vocabulary Separation and Adaptation
Jiong Wang, Zhou Zhao, Weike Jin +5
For face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The repre…