1 citations · 1 across the 4 of their papers we have counts for
4 papers
Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models
Benlei Cui, Shen Pang, Yuke Wang +7
The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based atta…
Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries
Xinfeng Li, Shengyuan Pang, Jialin Wu +5
Text-to-image (T2I) models, though exhibiting remarkable creativity in image generation, can be exploited to produce unsafe images. Existing safety measures, e.g., content moderati…
Protego: Detecting Adversarial Examples for Vision Transformers via Intrinsic Capabilities
Jialin Wu, Kaikai Pan, Yanjiao Chen +3
Transformer models have excelled in natural language tasks, prompting the vision community to explore their implementation in computer vision problems. However, these models are st…
Legilimens: Practical and Unified Content Moderation for Large Language Model Services
Jialin Wu, Jiangyi Deng, Shengyuan Pang +4
Given the societal impact of unsafe content generated by large language models (LLMs), ensuring that LLM services comply with safety standards is a crucial concern for LLM service…