4 papers
To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents
Wei Shi, Ziheng Peng, Sihang Li +4
LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high…
SAFER: Probing Safety in Reward Models with Sparse Autoencoder
Wei Shi, Ziyuan Xie, Sihang Li +1
Reinforcement learning from human feedback (RLHF) is a key paradigm for aligning large language models (LLMs) with human values, yet the reward models at its core remain largely op…
Interpretable Reward Model via Sparse Autoencoder
Shuyi Zhang, Wei Shi, Sihang Li +3
Large language models (LLMs) have been widely deployed across numerous fields. Reinforcement Learning from Human Feedback (RLHF) leverages reward models (RMs) as proxies for human…
Route Sparse Autoencoder to Interpret Large Language Models
Wei Shi, Sihang Li, Tao Liang +4
Mechanistic interpretability of large language models (LLMs) aims to uncover the internal processes of information propagation and reasoning. Sparse autoencoders (SAEs) have demons…