4 citations · 7 across the 5 of their papers we have counts for
1 paper · 1 filter
Ang Li, Qiugen Xiao, Peng Cao +12
Reinforcement Learning from AI Feedback (RLAIF) has the advantages of shorter annotation cycles and lower costs over Reinforcement Learning from Human Feedback (RLHF), making it hi…