1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Zachary Ankner, Mansheej Paul, Brandon Cui +2
Traditionally, reward models used for reinforcement learning from human feedback (RLHF) are trained to directly predict preference scores without leveraging the generation capabili…