14 citations · 14 across the 1 of their papers we have counts for
1 paper
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis +4
Large language models (LLMs) fine-tuned with reinforcement learning from human feedback (RLHF) have been used in some of the most widely deployed AI models to date, such as OpenAI'…