1 citations · 1 across the 1 of their papers we have counts for
1 paper
Jinghan Zhang, Xiting Wang, Yiqiao Jin +3
The reward model for Reinforcement Learning from Human Feedback (RLHF) has proven effective in fine-tuning Large Language Models (LLMs). Notably, collecting human feedback for RLHF…