4 citations · 7 across the 3 of their papers we have counts for
6 papers
HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models
Sensen Gao, Xiaojun Jia, Yihao Huang +5
Text-to-Image(T2I) models have achieved remarkable success in image generation and editing, yet these models still have many potential issues, particularly in generating inappropri…
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
Kuofeng Gao, Jindong Gu, Yang Bai +4
Despite the exceptional performance of multi-modal large language models (MLLMs), their deployment requires substantial computational resources. Once malicious users induce high en…
On the Multi-modal Vulnerability of Diffusion Models
Dingcheng Yang, Yang Bai, Xiaojun Jia +3
Diffusion models have been widely deployed in various image generation tasks, demonstrating an extraordinary connection between image and text modalities. Although prior studies ha…
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
Kuofeng Gao, Yang Bai, Jindong Gu +4
Large vision-language models (VLMs) such as GPT-4 have achieved exceptional performance across various multi-modal tasks. However, the deployment of VLMs necessitates substantial e…
OT-Attack: Enhancing Adversarial Transferability of Vision-Language Models via Optimal Transport Optimization
Dongchen Han, Xiaojun Jia, Yang Bai +3
Vision-language pre-training (VLP) models demonstrate impressive abilities in processing both images and text. However, they are vulnerable to multi-modal adversarial examples (AEs…
Fast Propagation is Better: Accelerating Single-Step Adversarial Training via Sampling Subnetworks
Xiaojun Jia, Jianshu Li, Jindong Gu +2
Adversarial training has shown promise in building robust models against adversarial examples. A major drawback of adversarial training is the computational overhead introduced by…