1 citations · 1 across the 4 of their papers we have counts for
8 papers
Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection
Zedian Shao, Hongbin Liu, Yuepeng Hu +1
Multi-modal large language models (MLLMs) have emerged as powerful tools for analyzing Internet-scale image data, offering significant benefits but also raising critical safety and…
VideoMarkBench: Benchmarking Robustness of Video Watermarking
Zhengyuan Jiang, Moyang Guo, Kecen Li +5
The rapid development of video generative models has led to a surge in highly realistic synthetic videos, raising ethical concerns related to disinformation and copyright infringem…
WebInject: Prompt Injection Attack to Web Agents
Xilong Wang, John Bloch, Zedian Shao +3
Multi-modal large language model (MLLM)-based web agents interact with webpage environments by generating actions based on screenshots of the webpages. In this work, we propose Web…
Zero-shot Autonomous Microscopy for Scalable and Intelligent Characterization of 2D Materials
Jingyun Yang, Ruoyan Avery Yin, Chi Jiang +14
Characterization of atomic-scale materials traditionally requires human experts with months to years of specialized training. Even for trained human operators, accurate and reliabl…
Jailbreaking Safeguarded Text-to-Image Models via Large Language Models
Zhengyuan Jiang, Yuepeng Hu, Yuchen Yang +2
Text-to-Image models may generate harmful content, such as pornographic images, particularly when unsafe prompts are submitted. To address this issue, safety filters are often adde…
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
Yuepeng Hu, Zhengyuan Jiang, Neil Zhenqiang Gong
Text-to-image models can generate harmful images when presented with unsafe prompts, posing significant safety and societal risks. Alignment methods aim to modify these models to e…