Patch is Enough: Naturalistic Adversarial Patch against Vision-Language Pre-training Models
arXiv:2410.04884 · doi:10.1007/s44267-024-00050-1
Abstract
Visual language pre-training (VLP) models have demonstrated significant success across various domains, yet they remain vulnerable to adversarial attacks. Addressing these adversarial vulnerabilities is crucial for enhancing security in multimodal learning. Traditionally, adversarial methods targeting VLP models involve simultaneously perturbing images and text. However, this approach faces notable challenges: first, adversarial perturbations often fail to translate effectively into real-world scenarios; second, direct modifications to the text are conspicuously visible. To overcome these limitations, we propose a novel strategy that exclusively employs image patches for attacks, thus preserving the integrity of the original text. Our method leverages prior knowledge from diffusion models to enhance the authenticity and naturalness of the perturbations. Moreover, to optimize patch placement and improve the efficacy of our attacks, we utilize the cross-attention mechanism, which encapsulates intermodal interactions by generating attention maps to guide strategic patch placements. Comprehensive experiments conducted in a white-box setting for image-to-text scenarios reveal that our proposed method significantly outperforms existing techniques, achieving a 100% attack success rate. Additionally, it demonstrates commendable performance in transfer tasks involving text-to-image configurations.
accepted by Visual Intelligence
References in corpus (17)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Concealed Object Detection
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day
- Fast Camouflaged Object Detection via Edge-based Reversible Re-calibration Network
- Structure-measure: A New Way to Evaluate Foreground Maps
- Light Field Salient Object Detection: A Review and Benchmark
- Language Models with Image Descriptors are Strong Few-Shot Video-Language Learners
- Salient Objects in Clutter: Bringing Salient Object Detection to the Foreground
- JL-DCF: Joint Learning and Densely-Cooperative Fusion Framework for RGB-D Salient Object Detection
- Siamese Network for RGB-D Salient Object Detection and Beyond
- How Good is Google Bard's Visual Understanding? An Empirical Study on Open Challenges
- SPot-the-Difference Self-Supervised Pre-training for Anomaly Detection and Segmentation
- AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models
- MOSE: A New Dataset for Video Object Segmentation in Complex Scenes
- Exposing and Mitigating Spurious Correlations for Cross-Modal Retrieval
- MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions