2 papers
cs.CL2025
Safety Alignment of Large Language Models via Contrasting Safe and Harmful Distributions
Xiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi +3
With the widespread application of Large Language Models (LLMs), it has become a significant concern to ensure their safety and prevent harmful responses. While current safe-alignm…
cs.CR2025
DiffZOO: A Purely Query-Based Black-Box Attack for Red-teaming Text-to-Image Generative Model via Zeroth Order Optimization
Pucheng Dang, Xing Hu, Dong Li +3
Current text-to-image (T2I) synthesis diffusion models raise misuse concerns, particularly in creating prohibited or not-safe-for-work (NSFW) images. To address this, various safet…