4 papers · 1 filter
Depth-Copy-Paste: Multimodal and Depth-Aware Compositing for Robust Face Detection
Qiushi Guo
Data augmentation is crucial for improving the robustness of face detection systems, especially under challenging conditions such as occlusion, illumination variation, and complex…
DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
Yifeng Gao, Yifan Ding, Hongyu Su +9
As AI-generated video becomes increasingly pervasive across media platforms, the ability to reliably distinguish synthetic content from authentic footage has become both urgent and…
Adversarial Prompt Distillation for Vision-Language Models
Lin Luo, Xin Wang, Bojia Zi +3
Large pre-trained Vision-Language Models (VLMs) such as Contrastive Language-Image Pre-training (CLIP) have been shown to be susceptible to adversarial attacks, raising concerns ab…
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
Yunhan Zhao, Xiang Zheng, Lin Luo +3
In this paper, we focus on black-box defense for VLMs against jailbreak attacks. Existing black-box defense methods are either unimodal or bimodal. Unimodal methods enhance either…