14 papers
When Muon Optimizer Meets Adversarial Training: A Theoretical and Empirical Study
Jun Yan, Weiquan Huang, Jiankai Zuo +4
Adversarial training (AT) remains one of the most reliable empirical defenses against adversarial attacks. Its robustness critically depends on how the underlying min-max objective…
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces
Chang Jin, An Wang, Zeming Wei +7
Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools, memory, and execution environm…
Securing Multi-Agent Systems Against Corruptions via Node Contribution Backpropagation
Chengcan Wu, Zhixin Zhang, Mingqian Xu +2
Multi-Agent Systems (MAS) have become a prevalent paradigm for Large Language Model (LLM) applications. However, the complex multi-agent design in MAS introduces unique trustworthi…
Symbolic-Neural Soft-Logic Reasoning: Towards Robust and Verifiable Thinking Chains via Cooperative Evolution
Rui Wang, Zeming Wei, Yihao Zhang +1
Large Language Models (LLMs) have demonstrated impressive progress in complex reasoning tasks, largely driven by the Chain-of-Thought (CoT) paradigm, which decomposes difficult pro…
RACC: Representation-Aware Coverage Criteria for LLM Safety Testing
Zeming Wei, Zhixin Zhang, Chengcan Wu +3
Large Language Models (LLMs) face severe safety risks from jailbreak attacks, yet current safety testing largely relies on static datasets and lacks systematic criteria to evaluate…
Secure LLM Fine-Tuning via Safety-Aware Probing
Chengcan Wu, Zhixin Zhang, Zeming Wei +3
Large language models (LLMs) have achieved remarkable success across many applications, but their ability to generate harmful content raises serious safety concerns. Although safet…