1 paper · 1 filter
Abdulhady Abas Abdullah, Fatemeh Daneshfar, Seyedali Mirjalili +1
Aligning large language models (LLMs) with human preferences is commonly done via reinforcement learning from human feedback (RLHF) with Proximal Policy Optimization (PPO) or, more…