Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
Xiaomin Li, Xupeng Chen, Jingxuan Fan +2
The safety alignment of large language models (LLMs) often relies on reinforcement learning from human feedback (RLHF), which requires human annotations to construct preference dat…
cs.CL2025
CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
Sijia Chen, Xiaomin Li, Mengxue Zhang +3
Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While…