1 paper
Zhiyu Mei, Wei Fu, Kaiwei Li +3
Reinforcement Learning from Human Feedback (RLHF) is a pivotal technique for empowering large language model (LLM) applications. Compared with the supervised training process of LL…