1 citations · 1 across the 1 of their papers we have counts for
1 paper
Yingshui Tan, Yilei Jiang, Yanshi Li +6
Fine-tuning large language models (LLMs) based on human preferences, commonly achieved through reinforcement learning from human feedback (RLHF), has been effective in improving th…