3 papers
cs.CL2026
A Few Neurons Reveal When LLMs Misuse Tools: Sparse Detection and Selective Steering for Reliable Tool Use
Yutong Ke, Ming Yin, Chongwen Zhao +1
Agentic LLMs exhibit three consequential tool-use failures: invalid arguments (validity), unnecessary calls (over-calling), and omitted calls when tools are needed (missing). We fi…
cs.AI2026
Unraveling LLM Jailbreaks Through Safety Knowledge Neurons
Chongwen Zhao, Yutong Ke, Kaizhu Huang
Large Language Models (LLMs) are increasingly attracting attention in various applications. Nonetheless, there is a growing concern as some users attempt to exploit these models fo…
cs.AI2025
Defending against Jailbreak through Early Exit Generation of Large Language Models
Chongwen Zhao, Zhihao Dou, Kaizhu Huang
Large Language Models (LLMs) are increasingly attracting attention in various applications. Nonetheless, there is a growing concern as some users attempt to exploit these models fo…