1 citations · 1 across the 1 of their papers we have counts for
1 paper
Feiteng Fang, Liang Zhu, Min Yang +6
Reinforcement learning from human feedback (RLHF) is a crucial technique in aligning large language models (LLMs) with human preferences, ensuring these LLMs behave in beneficial a…