6 papers
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
Lingzhi Yuan, Xinfeng Li, Chejian Xu +6
Recent text-to-image (T2I) models have exhibited remarkable performance in generating high-quality images from text descriptions. However, these models are vulnerable to misuse, pa…
A Sentence Relation-Based Approach to Sanitizing Malicious Instructions
Soumil Datta, Melissa Umble, Daniel S. Brown +1
Retrieval-augmented generation and tool-integrated LLM agents increasingly depend on external textual sources. This reliance broadens the available attack surface, allowing adversa…
How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies
Akansha Kalra, Basavasagar Patil, Guanhong Tao +1
Learning from demonstrations is a popular approach to train AI models; however, their vulnerability to adversarial attacks remains underexplored. We present the first systematic st…
Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models
Xiaomei Zhang, Zhaoxi Zhang, Leo Yu Zhang +3
Visual token compression is widely adopted to improve the inference efficiency of Large Vision-Language Models (LVLMs), enabling their deployment in latency-sensitive and resource-…
Dataset Poisoning Attacks on Behavioral Cloning Policies
Akansha Kalra, Soumil Datta, Ethan Gilmore +3
Behavior Cloning (BC) is a popular framework for training sequential decision policies from expert demonstrations via supervised learning. As these policies are increasingly being…
Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
Zhiyuan Zhong, Zhen Sun, Yepang Liu +2
Vision Language Models (VLMs) have shown remarkable performance, but are also vulnerable to backdoor attacks whereby the adversary can manipulate the model's outputs through hidden…