5 papers
Attention Sinks and Outliers in Attention Residuals
Haozheng Luo, Haoran Dai, Shaoyang Zhang +10
We propose OASIS, an outlier- and sink-aware technique built on inter-layer null signaling. As AttnResidual architectures introduce an additional depth-wise normalization channel,…
You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents
Ching-Yu Kao, Xinfeng Li, Shenyu Dai +4
High-privilege LLM agents that autonomously process external documentation are increasingly trusted to automate tasks by reading and executing project instructions, yet they are gr…
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
Xiaomin Li, Xupeng Chen, Jingxuan Fan +2
The safety alignment of large language models (LLMs) often relies on reinforcement learning from human feedback (RLHF), which requires human annotations to construct preference dat…
Learning to Rank Chain-of-Thought: Using a Small Model
Eric Hanchen Jiang, Haozheng Luo, Shengyuan Pang +9
Large Language Models (LLMs) struggle with reliable mathematical reasoning, and current verification methods are often computationally expensive. This paper introduces the Energy O…
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
Sijia Chen, Xiaomin Li, Mengxue Zhang +3
Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While…