1 paper
Prakhar Gupta, Vaibhav Gupta
Post-training with reinforcement learning (RL) typically optimizes a single scalar objective and ignores structure in how solutions are produced. We ask whether a scalar hint towar…