1 paper
Hongyu Yang, Qi Zhao, Zhenhua hu +1
Reinforcement Learning from Human Feedback and its variants excel in aligning with human intentions to generate helpful, harmless, and honest responses. However, most of them rely…