1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Peng Kuang, Yanli Wang, Xiaoyu Han +3
Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language models (LLMs). However, this promise…