1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Yulan Hu, Qingyang Li, Sheng Ouyang +6
Reinforcement Learning from Human Feedback (RLHF) facilitates the alignment of large language models (LLMs) with human preferences, thereby enhancing the quality of responses gener…