1 paper · 1 filter
Zhiyu Mei, Wei Fu, Kaiwei Li +3
Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for empowering large language model (LLM) applications. Compared with the supervised training process of LL…