3 citations · 3 across the 1 of their papers we have counts for
1 paper · 1 filter
Nathan Lambert
Reinforcement learning from human feedback (RLHF) has become a crucial tool to build the latest machine learning systems at scale. The field grew around the core methods of RLHF in…