4 papers
DTVI: Dual-Stage Textual and Visual Intervention for Safe Text-to-Image Generation
Binhong Tan, Zhaoxin Wang, Handing Wang
Text-to-Image (T2I) diffusion models have demonstrated strong generation ability, but their potential to generate unsafe content raises significant safety concerns. Existing infere…
Multilingual Safety Alignment Via Sparse Weight Editing
Jiaming Liang, Zhaoxin Wang, Handing Wang
Large Language Models (LLMs) exhibit significant safety disparities across languages, with low-resource languages (LRLs) often bypassing safety guardrails established for high-reso…
SafeNeuron: Neuron-Level Safety Alignment for Large Language Models
Zhaoxin Wang, Jiaming Liang, Fengbin Zhu +5
Large language models (LLMs) and multimodal LLMs are typically safety-aligned before release to prevent harmful content generation. However, recent studies show that safety behavio…
From Parameter to Representation: A Closed-Form Approach for Controllable Model Merging
Jialin Wu, Jian Yang, Handing Wang +2
Model merging combines expert models for multitask performance but faces challenges from parameter interference. This has sparked recent interest in controllable model merging, giv…