1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yuanzhao Zhai, Han Zhang, Yu Lei +5
Reinforcement learning from human feedback (RLHF) emerges as a promising paradigm for aligning large language models (LLMs). However, a notable challenge in RLHF is overoptimizatio…