4 citations · 4 across the 11 of their papers we have counts for
4 papers · 1 filter
RARE: Decoupling Representation Steering from Expert Routing in Mixture-of-Experts Language Models
Zhibo Zhang, Zhen Ouyang, Ling Shi +1
Representation engineering offers a lightweight means of controlling language-model behavior by modifying intermediate hidden states, but its direct application to Mixture-of-Exper…
RASET: Router-Agnostic Safety-Critical Expert Tuning Exposes Localized Safety Enforcement Failures in Mixture-of-Experts LLMs
Zhibo Zhang, Yuxi Li, Zhen Ouyang +2
Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization remains underexplored. A common…
Detecting LLM Fact-conflicting Hallucinations Enhanced by Temporal-logic-based Reasoning
Ningke Li, Yahui Song, Kailong Wang +4
Large language models (LLMs) face the challenge of hallucinations -- outputs that seem coherent but are actually incorrect. A particularly damaging type is fact-conflicting halluci…
Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective Detection
Yuxi Li, Yi Liu, Gelei Deng +7
With the expanding application of Large Language Models (LLMs) in various domains, it becomes imperative to comprehensively investigate their unforeseen behaviors and consequent ou…