4 papers · 1 filter
RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs
Zhibo Zhang, Yuxi Li, Zhen Ouyang +2
Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization remains underexplored. A common…
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
Zhibo Zhang, Yuxi Li, Kailong Wang +3
Large Language Models (LLMs) have achieved remarkable success across domains such as healthcare, education, and cybersecurity. However, this openness also introduces significant se…
GlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models
Zhibo Zhang, Wuxia Bai, Yuxi Li +6
Large language models (LLMs) have achieved unprecedented success in the field of natural language processing. However, the black-box nature of their internal mechanisms has brought…
Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective Detection
Yuxi Li, Yi Liu, Gelei Deng +7
With the expanding application of Large Language Models (LLMs) in various domains, it becomes imperative to comprehensively investigate their unforeseen behaviors and consequent ou…