1 citations · 1 across the 8 of their papers we have counts for
1 paper · 2 filters
Junkang Wu, Xue Wang, Zhengyi Yang +5
Aligning large language models (LLMs) with human values and intentions is crucial for their utility, honesty, and safety. Reinforcement learning from human feedback (RLHF) is a pop…