4 papers
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
Yingxin Lai, Zitong Yu, Jun Wang +3
Multimodal Large Language Models (MLLMs) enable interpretable multimedia forensics by generating textual rationales for forgery detection. However, processing dense visual sequence…
DADM: Dual Alignment of Domain and Modality for Face Anti-spoofing
Jingyi Yang, Xun Lin, Zitong Yu +5
With the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a promi…
GVformer: Graph Guided Video Vision Transformer for Face Anti-Spoofing
Jingyi Yang, Zitong Yu, Xiuming Ni +2
In videos containing spoofed faces, we may uncover the spoofing evidence based on either photometric or dynamic abnormality, even a combination of both. Prevailing face anti-spoofi…
Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
Jingyi Yang, Zitong Yu, Xiuming Ni +2
Contrastive language-image pretraining (CLIP) has significantly advanced image-based vision learning. A pressing topic subsequently arises: how can we effectively adapt CLIP to the…