2 citations · 2 across the 1 of their papers we have counts for
1 paper
Di Jin, Shikib Mehri, Devamanyu Hazarika +4
Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF)…