activity
20242026
most citedGlitchProber: Advancing Effective Detection and Mitigation of Glitch Tokens in Large Language Models

2 citations · 2 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CL2026

RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models

Zhibo Zhang, Zhen Ouyang, Ling Shi +1

Representation engineering offers a lightweight means of controlling language-model behavior by modifying intermediate hidden states, but its direct application to Mixture-of-Exper…

cs.AI2026

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

Zhibo Zhang, Zhen Ouyang, Ling Shi +1

Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automa…

cs.CR2026

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning

Yuxi Li, Zhibo Zhang, Kailong Wang +3

Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or t…

cs.CL2026

RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs

Zhibo Zhang, Yuxi Li, Zhen Ouyang +2

Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization remains underexplored. A common…

cs.CR2026

When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion

Jiaqing Li, Zhibo Zhang, Shide Zhou +3

Model merging has emerged as a powerful technique for combining specialized capabilities from multiple fine-tuned LLMs without additional training costs. However, the security impl…

cs.CR2025

Embedding Poisoning: Bypassing Safety Alignment via Embedding Semantic Shift

Shuai Yuan, Zhibo Zhang, Yuxi Li +2

The widespread distribution of Large Language Models (LLMs) through public platforms like Hugging Face introduces significant security challenges. While these platforms perform bas…