4 papers
Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training
Mengnan Zhao, Geyong Min, Lihe Zhang +2
Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward head classes, and the adversarial inner…
Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models
Weiwei Qi, Zefeng Wu, Tianhang Zheng +4
Ensuring Large Language Model (LLM) safety is crucial, yet the lack of a clear understanding about safety mechanisms hinders the development of precise and reliable methodologies f…
AdvAnchor: Enhancing Diffusion Model Unlearning with Adversarial Anchors
Mengnan Zhao, Lihe Zhang, Xingyi Yang +2
Security concerns surrounding text-to-image diffusion models have driven researchers to unlearn inappropriate concepts through fine-tuning. Recent fine-tuning methods typically ali…
Separable Multi-Concept Erasure from Diffusion Models
Mengnan Zhao, Lihe Zhang, Tianhang Zheng +2
Large-scale diffusion models, known for their impressive image generation capabilities, have raised concerns among researchers regarding social impacts, such as the imitation of co…