1 paper · 1 filter
Jie Cheng, Gang Xiong, Ruixi Qiao +5
Process reward models (PRMs) have proven effective for test-time scaling of Large Language Models (LLMs) on challenging reasoning tasks. However, reward hacking issues with PRMs li…