1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Ruijie Xu, Zhihan Liu, Yongfei Liu +4
We address the challenge of online Reinforcement Learning from Human Feedback (RLHF) with a focus on self-rewarding alignment methods. In online RLHF, obtaining feedback requires i…