6 papers
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
Yuxi Li, Zhibo Zhang, Kailong Wang +3
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or t…
Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs
Zhibo Zhang, Yuxi Li, Zhen Ouyang +2
Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization remains underexplored. A common…
When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion
Jiaqing Li, Zhibo Zhang, Shide Zhou +3
Model merging has emerged as a powerful technique for combining specialized capabilities from multiple fine-tuned LLMs without additional training costs. However, the security impl…
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
Shuai Yuan, Zhibo Zhang, Yuxi Li +2
The widespread distribution of Large Language Models (LLMs) through public platforms like Hugging Face introduces significant security challenges. While these platforms perform bas…
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
Zhibo Zhang, Yuxi Li, Kailong Wang +3
Large Language Models (LLMs) have achieved remarkable success across domains such as healthcare, education, and cybersecurity. However, this openness also introduces significant se…
RefPentester: A Knowledge-Informed Self-Reflective Penetration Testing Framework Based on Large Language Models
Hanzheng Dai, Yuanliang Li, Jun Yan +1
Automated penetration testing (AutoPT) powered by large language models (LLMs) has gained attention for its ability to automate ethical hacking processes and identify vulnerabiliti…