activity
20232026
collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG2026

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

Hongli Shen, Shaopeng Fu, Qinbo Zhang +2

Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outputs. Recent methods align LRMs using direc…

cs.LG2026

Benign Overfitting in Adversarial Training for Vision Transformers

Jiaming Zhang, Meng Ding, Shaopeng Fu +2

Despite the remarkable success of Vision Transformers (ViTs) across a wide range of vision tasks, recent studies have revealed that they remain vulnerable to adversarial examples,…

cs.LG2026

Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory

Shaopeng Fu, Di Wang

Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To improve the efficiency of AT f…

cs.LG2026

Understanding the Impact of Differentially Private Training on Memorization of Long-Tailed Data

Jiaming Zhang, Huanyi Xie, Meng Ding +3

Recent research shows that modern deep learning models achieve high predictive accuracy partly by memorizing individual training samples. Such memorization raises serious privacy c…

cs.LG2025

Understanding Private Learning From Feature Perspective

Meng Ding, Mingxi Lei, Shaopeng Fu +3

Differentially private Stochastic Gradient Descent (DP-SGD) has become integral to privacy-preserving machine learning, ensuring robust privacy guarantees in sensitive domains. Des…

cs.LG2025

Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence

Shaopeng Fu, Liang Ding, Jingfeng Zhang +1

Jailbreak attacks against large language models (LLMs) aim to induce harmful behaviors in LLMs through carefully crafted adversarial prompts. To mitigate attacks, one way is to per…