6 citations · 20 across the 22 of their papers we have counts for
1 paper · 1 filter
Yu Zhu, Chuxiong Sun, Wenfei Yang +8
Reinforcement Learning from Human Feedback (RLHF) is the prevailing approach to ensure Large Language Models (LLMs) align with human values. However, existing RLHF methods require…