2 citations · 2 across the 3 of their papers we have counts for
4 papers
Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift
Shuai Yuan, Zhibo Zhang, Yuxi Li +2
The widespread distribution of Large Language Models (LLMs) through public platforms like Hugging Face introduces significant security challenges. While these platforms perform bas…
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
Zhibo Zhang, Yuxi Li, Kailong Wang +3
Large Language Models (LLMs) have achieved remarkable success across domains such as healthcare, education, and cybersecurity. However, this openness also introduces significant se…
Model-Editing-Based Jailbreak against Safety-aligned Large Language Models
Yuxi Li, Zhibo Zhang, Kailong Wang +2
Large Language Models (LLMs) have transformed numerous fields by enabling advanced natural language interactions but remain susceptible to critical vulnerabilities, particularly ja…
GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
Zhibo Zhang, Wuxia Bai, Yuxi Li +6
Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the black-box nature of their internal mechanisms has brought…