3 citations · 9 across the 5 of their papers we have counts for
1 paper · 1 filter
Wei Shen, Xiaoying Zhang, Yuanshun Yao +3
Reinforcement learning from human feedback (RLHF) is the mainstream paradigm used to align large language models (LLMs) with human preferences. Yet existing RLHF heavily relies on…