Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents
Wei Shi, Ziheng Peng, Sihang Li +4
LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high…
cs.LG2026
Differentially Private Subspace Fine-Tuning for Large Language Models
Lele Zheng, Xiang Wang, Tao Zhang +3
Fine-tuning large language models on downstream tasks is crucial for realizing their cross-domain potential but often relies on sensitive data, raising privacy concerns. Differenti…
cs.LG2025
Interpretable Reward Model via Sparse Autoencoder
Shuyi Zhang, Wei Shi, Sihang Li +3
Large language models (LLMs) have been widely deployed across numerous fields. Reinforcement Learning from Human Feedback (RLHF) leverages reward models (RMs) as proxies for human…