1 paper · 1 filter
Zhilin Wang, Alexander Bukharin, Olivier Delalleau +5
Reward models are critical for aligning models to follow instructions, and are typically trained following one of two popular paradigms: Bradley-Terry style or Regression style. Ho…