most citedInducing High Energy-Latency of Large Vision-Language Models with Verbose Images

4 citations · 7 across the 3 of their papers we have counts for

collaborators

6 papers

cs.CV2024

HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models

Sensen Gao, Xiaojun Jia, Yihao Huang +5

Text-to-Image(T2I) models have achieved remarkable success in image generation and editing, yet these models still have many potential issues, particularly in generating inappropri…

cs.CV20242 cited

Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples

Kuofeng Gao, Jindong Gu, Yang Bai +4

Despite the exceptional performance of multi-modal large language models (MLLMs), their deployment requires substantial computational resources. Once malicious users induce high en…

cs.LG2024

On the Multi-modal Vulnerability of Diffusion Models

Dingcheng Yang, Yang Bai, Xiaojun Jia +3

Diffusion models have been widely deployed in various image generation tasks, demonstrating an extraordinary connection between image and text modalities. Although prior studies ha…

cs.CV20244 cited

Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images

Kuofeng Gao, Yang Bai, Jindong Gu +4

Large vision-language models (VLMs) such as GPT-4 have achieved exceptional performance across various multi-modal tasks. However, the deployment of VLMs necessitates substantial e…

cs.CV2023

OT-Attack: Enhancing Adversarial Transferability of Vision-Language Models via Optimal Transport Optimization

Dongchen Han, Xiaojun Jia, Yang Bai +3

Vision-language pre-training (VLP) models demonstrate impressive abilities in processing both images and text. However, they are vulnerable to multi-modal adversarial examples (AEs…

cs.CV20231 cited

Fast Propagation is Better: Accelerating Single-Step Adversarial Training via Sampling Subnetworks

Xiaojun Jia, Jianshu Li, Jindong Gu +2

Adversarial training has shown promise in building robust models against adversarial examples. A major drawback of adversarial training is the computational overhead introduced by…