2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Simeng Sun, Dhawal Gupta, Mohit Iyyer
During the last stage of RLHF, a large language model is aligned to human intents via PPO training, a process that generally requires large-scale computational resources. In this t…