4 papers
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
Xiaomin Li, Xupeng Chen, Jingxuan Fan +2
The safety alignment of large language models (LLMs) often relies on reinforcement learning from human feedback (RLHF), which requires human annotations to construct preference dat…
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
Xiaomin Li, Mingye Gao, Zhiwei Zhang +2
High-quality training data is critical to the performance of large language models (LLMs). Recent work has explored using LLMs to rate and select data based on a small set of human…
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Xiaomin Li, Mingye Gao, Yuexing Hao +4
Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making.…
Data-adaptive Safety Rules for Training Reward Models
Xiaomin Li, Mingye Gao, Zhiwei Zhang +2
Reinforcement Learning from Human Feedback (RLHF) is commonly employed to tailor models to human preferences, especially to improve the safety of outputs from large language models…