1 paper
Rudransh Agnihotri, Ananya Pandey
Reward-model training is the cost bottleneck in modern Reinforcement Learning Human Feedback (RLHF) pipelines, often requiring tens of billions of parameters and an offline prefere…