1 paper · 2 filters
Yujie Zhao, Jose Efraim Aguilar Escamill, Weyl Lu +1
Reinforcement Learning from Human Feedback (RLHF) has recently surged in popularity, particularly for aligning large language models and other AI systems with human intentions. At…