activity
20242026
collaborators
Showing cs.LGShow all

9 papers · 1 filter

cs.LG2026

When Muon Optimizer Meets Adversarial Training: A Theoretical and Empirical Study

Jun Yan, Weiquan Huang, Jiankai Zuo +4

Adversarial training (AT) remains one of the most reliable empirical defenses against adversarial attacks. Its robustness critically depends on how the underlying min-max objective…

cs.LG2026

Secure LLM Fine-Tuning via Safety-Aware Probing

Chengcan Wu, Zhixin Zhang, Zeming Wei +3

Large language models (LLMs) have achieved remarkable success across many applications, but their ability to generate harmful content raises serious safety concerns. Although safet…

cs.LG2026

Absorber LLM: Harnessing Causal Synchronization for Test-Time Training

Zhixin Zhang, Shabo Zhang, Chengcan Wu +2

Transformers suffer from a high computational cost that grows with sequence length for self-attention, making inference in long streams prohibited by memory consumption. Constant-m…

cs.LG2025

Stabilizing Multi-Attack Adversarial Training via Bandit Optimization

Rui Wang, Zeming Wei, Xiyue Zhang +1

Deep Neural Networks (DNNs) remain vulnerable to diverse adversarial perturbations, motivating multi-attack adversarial training (AT) for improved robustness. However, existing met…

cs.LG2025

Dynamic Orthogonal Continual Fine-tuning for Mitigating Catastrophic Forgettings

Zhixin Zhang, Zeming Wei, Meng Sun

Catastrophic forgetting remains a critical challenge in continual learning for large language models (LLMs), where models struggle to retain performance on historical tasks when fi…

cs.LG2025

Boosting Jailbreak Attack with Momentum

Yihao Zhang, Zeming Wei

Large Language Models (LLMs) have achieved remarkable success across diverse tasks, yet they remain vulnerable to adversarial attacks, notably the well-known jailbreak attack. In p…