1 paper
Yuanzhao Zhai, Zhuo Zhang, Kele Xu +6
Aligning with human preference datasets has been critical to the success of large language models (LLMs). Reinforcement learning from human feedback (RLHF) employs a costly reward…