1 paper
Yangyang Zhou, Yi-Chen Li
Reinforcement Learning from Human Feedback has become the standard paradigm for language model alignment, where reward models directly determine alignment effectiveness. In this wo…