1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Yankai Yang, Yancheng Long, Hongyang Wei +12
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of generative models. For complex tasks such as i…