2 citations · 2 across the 1 of their papers we have counts for
1 paper
Nathan Lambert, Thomas Krendl Gilbert, Tom Zick
Reinforcement learning from human feedback (RLHF) has emerged as a powerful technique to make large language models (LLMs) easier to use and more effective. A core piece of the RLH…