4 papers
Understanding the Impact of Differentially Private Training on Memorization of Long-Tailed Data
Jiaming Zhang, Huanyi Xie, Meng Ding +3
Recent research shows that modern deep learning models achieve high predictive accuracy partly by memorizing individual training samples. Such memorization raises serious privacy c…
Understanding Private Learning From Feature Perspective
Meng Ding, Mingxi Lei, Shaopeng Fu +3
Differentially private Stochastic Gradient Descent (DP-SGD) has become integral to privacy-preserving machine learning, ensuring robust privacy guarantees in sensitive domains. Des…
Backdooring CLIP through Concept Confusion
Lijie Hu, Junchi Liao, Weimin Lyu +5
Backdoor attacks pose a serious threat to deep learning models by allowing adversaries to implant hidden behaviors that remain dormant on clean inputs but are maliciously triggered…
Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence
Shaopeng Fu, Liang Ding, Jingfeng Zhang +1
Jailbreak attacks against large language models (LLMs) aim to induce harmful behaviors in LLMs through carefully crafted adversarial prompts. To mitigate attacks, one way is to per…