1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Saeed Khaki, JinJin Li, Lan Ma +2
Reinforcement learning from human feedback (RLHF) has been extensively employed to align large language models with user intent. However, proximal policy optimization (PPO) based R…