7 papers
Metaphor-based Jailbreak Attacks on Text-to-Image Models
Chenyu Zhang, Lanjun Wang, Yiwen Ma +3
Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversaria…
T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model
Chenyu Zhang, Tairen Zhang, Lanjun Wang +3
Using risky text prompts, such as pornography and violent prompts, to test the safety of text-to-image (T2I) models is a critical task. However, existing risky prompt datasets are…
Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning
Chenyu Zhang, Lanjun Wang, Yiwen Ma +2
Text-to-Image(T2I) models typically deploy safety filters to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructi…
Domain Adaptation from Generated Multi-Weather Images for Unsupervised Maritime Object Classification
Dan Song, Shumeng Huo, Wenhui Li +3
The classification and recognition of maritime objects are crucial for enhancing maritime safety, monitoring, and intelligent sea environment prediction. However, existing unsuperv…
MarkPlugger: Generalizable Watermark Framework for Latent Diffusion Models without Retraining
Guokai Zhang, Lanjun Wang, Yuting Su +1
Today, the family of latent diffusion models (LDMs) has gained prominence for its high quality outputs and scalability. This has also raised security concerns on social media, as m…
TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models
Ruidong Chen, Honglin Guo, Lanjun Wang +3
Recent advances in text-to-image diffusion models enable photorealistic image generation, but they also risk producing malicious content, such as NSFW images. To mitigate risk, con…