1 paper · 1 filter
Hongyu Yang, Qi Zhao, Zhenhua hu +1
Reinforcement Learning from Human Feedback and its variants excel in aligning with human intentions to generate helpful, harmless, and honest responses. However, most of them rely…