1 paper
Xiaomin Li, Mingye Gao, Zhiwei Zhang +2
Reinforcement Learning from Human Feedback (RLHF) is commonly employed to tailor models to human preferences, especially to improve the safety of outputs from large language models…