14 papers
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
Yuxi Li, Zhibo Zhang, Kailong Wang +3
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or t…
A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation
Shide Zhou, Kailong Wang, Ling Shi +1
Defining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge. Current benchmarks primarily rely on static datasets s…
R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
Aijia Cheng, Kailong Wang, Ling Shi +1
Function calling empowers large language models (LLMs) to interface with external tools, yet existing RL-based approaches suffer from misalignment between reasoning processes and t…
Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs
Zhibo Zhang, Yuxi Li, Zhen Ouyang +2
Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization remains underexplored. A common…
Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics
Shide Zhou, Kailong Wang, Ling Shi +1
The widespread adoption of Large Language Models (LLMs) in critical applications has introduced severe reliability and security risks, as LLMs remain vulnerable to notorious threat…
STEAMROLLER: A Multi-Agent System for Inclusive Automatic Speech Recognition for People who Stutter
Ziqi Xu, Yi Liu, Yuekang Li +3
People who stutter (PWS) face systemic exclusion in today's voice-driven society, where access to voice assistants, authentication systems, and remote work tools increasingly depen…