8 citations · 8 across the 1 of their papers we have counts for
1 paper
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe +3
Reinforcement learning from human feedback (RLHF) is fundamentally limited by the capacity of humans to correctly evaluate model output. To improve human evaluation ability and ove…