1 paper
Amrith Setlur, Zijian Wang, Andrew Cohen +2
Typical reinforcement learning (RL) methods for LLM reasoning waste compute on hard problems, where correct on-policy traces are rare, policy gradients vanish, and learning stalls.…