8 citations · 8 across the 12 of their papers we have counts for
1 paper · 1 filter
Xiaomin Li, Mingye Gao, Zhiwei Zhang +2
Reinforcement Learning from Human Feedback (RLHF) is commonly employed to tailor models to human preferences, especially to improve the safety of outputs from large language models…