20 papers
ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration
Bei Yan, Yuecong Min, Jie Zhang +2
Large Vision-Language Models (LVLMs) frequently suffer from severe hallucination issues. Existing mitigation strategies predominantly rely on isolated, single-step states to enhanc…
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
Junqi Yang, Yuecong Min, Jie Zhang +2
Despite rapid progress, Video Large Language Models (Video-LLMs) remain unreliable due to hallucinations, which are outputs that contradict either video evidence (faithfulness) or…
What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models
Sen Nie, Jie Zhang, Zhongqi Wang +3
Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and challenging trade-off. In this work,…
Steering Vision-Language Pre-trained Models for Incremental Face Presentation Attack Detection
Haoze Li, Jie Zhang, Guoying Zhao +2
Face Presentation Attack Detection (PAD) demands incremental learning (IL) to combat evolving spoofing tactics and domains. Privacy regulations, however, forbid retaining past data…
Dual Attention Guided Defense Against Malicious Edits
Jie Zhang, Shuai Dong, Shiguang Shan +1
Recent progress in text-to-image diffusion models has transformed image editing via text prompts, yet this also introduces significant ethical challenges from potential misuse in c…
Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
Shuai Dong, Jie Zhang, Guoying Zhao +2
Text-guided image editing via diffusion models, while powerful, raises significant concerns about misuse, motivating efforts to immunize images against unauthorized edits using imp…