6 papers
DTVI: Dual-Stage Textual and Visual Intervention for Safe Text-to-Image Generation
Binhong Tan, Zhaoxin Wang, Handing Wang
Text-to-Image (T2I) diffusion models have demonstrated strong generation ability, but their potential to generate unsafe content raises significant safety concerns. Existing infere…
Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models
Jiangtao Liu, Zhaoxin Wang, Handing Wang +2
Text-to-Image (T2I) generation has advanced rapidly in recent years, but they also raise safety concerns due to the potential production of harmful content. In the practical deploy…
Multilingual Safety Alignment Via Sparse Weight Editing
Jiaming Liang, Zhaoxin Wang, Handing Wang
Large Language Models (LLMs) exhibit significant safety disparities across languages, with low-resource languages (LRLs) often bypassing safety guardrails established for high-reso…
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
Zhaoxin Wang, Jiaming Liang, Fengbin Zhu +5
Large language models (LLMs) and multimodal LLMs are typically safety-aligned before release to prevent harmful content generation. However, recent studies show that safety behavio…
Enhancing the Effectiveness and Durability of Backdoor Attacks in Federated Learning through Maximizing Task Distinction
Zhaoxin Wang, Handing Wang, Cong Tian +1
Federated learning allows multiple participants to collaboratively train a central model without sharing their private data. However, this distributed nature also exposes new attac…
Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
Zhaoxin Wang, Handing Wang, Cong Tian +1
Multimodal large language models (MLLMs) enable powerful cross-modal reasoning capabilities. However, the expanded input space introduces new attack surfaces. Previous jailbreak at…