1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Tianze Wang, Dongnan Gui, Yifan Hu +2
Reinforcement Learning from Human Feedback (RLHF) has shown promise in aligning large language models (LLMs). Yet its reliance on a singular reward model often overlooks the divers…