2 citations · 2 across the 1 of their papers we have counts for
1 paper
Shengyi Huang, Michael Noukhovitch, Arian Hosseini +3
This work is the first to openly reproduce the Reinforcement Learning from Human Feedback (RLHF) scaling behaviors reported in OpenAI's seminal TL;DR summarization work. We create…