3 papers
cs.CV2026
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
Yingxin Lai, Zitong Yu, Jun Wang +3
Multimodal Large Language Models (MLLMs) enable interpretable multimedia forensics by generating textual rationales for forgery detection. However, processing dense visual sequence…
cs.CV2025
DADM: Dual Alignment of Domain and Modality for Face Anti-spoofing
Jingyi Yang, Xun Lin, Zitong Yu +5
With the availability of diverse sensor modalities (i.e., RGB, Depth, Infrared) and the success of multi-modal learning, multi-modal face anti-spoofing (FAS) has emerged as a promi…
cs.CV2025
Kronecker Mask and Interpretive Prompts are Language-Action Video Learners
Jingyi Yang, Zitong Yu, Xiuming Ni +2
Contrastive language-image pretraining (CLIP) has significantly advanced image-based vision learning. A pressing topic subsequently arises: how can we effectively adapt CLIP to the…