2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Sriyash Poddar, Yanming Wan, Hamish Ivison +2
Reinforcement Learning from Human Feedback (RLHF) is a powerful paradigm for aligning foundation models to human values and preferences. However, current RLHF techniques cannot acc…