1 paper
Aneesh Pappu, Billy Porter, Ilia Shumailov +1
Reinforcement learning with human feedback (RLHF) has become the dominant method to align large models to user preferences. Unlike fine-tuning, for which there are many studies reg…