11 papers
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
Jiayu Zhang, Shuo Ye, Qilang Ye +3
Audio-Visual Question Answering (AVQA) requires models to effectively utilize both visual and auditory modalities to answer complex and diverse questions about audio-visual scenes.…
Purify then Guide: Rethinking Domain Generalization for Multimodal Face Anti-Spoofing
Yingjie Ma, Xun Lin, Zitong Yu +7
Face Anti-Spoofing (FAS) is essential for the security of facial recognition systems in diverse scenarios such as payment processing and surveillance. Current multimodal FAS method…
SVC 2026: the Second Multimodal Deception Detection Challenge and the First Domain Generalized Remote Physiological Measurement Challenge
Dongliang Zhu, Zhiyi Niu, Bo Zhao +14
Subtle visual signals, although difficult to perceive with the naked eye, contain important information that can reveal hidden patterns in visual data. These signals play a key rol…
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
Yingjie Ma, Xun Lin, Yong Xu +2
Face anti-spoofing (FAS) has recently advanced in multimodal fusion, cross-domain generalization, and interpretability. With large language models and reinforcement learning (RL),…
FaceShield: Explainable Face Anti-Spoofing with Multimodal Large Language Models
Hongyang Wang, Yichen Shi, Zhuofu Tao +7
Face anti-spoofing (FAS) is crucial for protecting facial recognition systems from presentation attacks. Previous methods approached this task as a classification problem, lacking…
SVC 2025: the First Multimodal Deception Detection Challenge
Xun Lin, Xiaobao Guo, Taorui Wang +5
Deception detection is a critical task in real-world applications such as security screening, fraud prevention, and credibility assessment. While deep learning methods have shown p…