1 paper
Jie Cheng, Gang Xiong, Ruixi Qiao +5
Process reward models (PRMs) have proven effective for test-time scaling of Large Language Models (LLMs) on challenging reasoning tasks. However, reward hacking issues with PRMs li…