1 citations · 1 across the 1 of their papers we have counts for
1 paper
Hanyin Wang, Chufan Gao, Qiping Xu +9
Process-supervised reward models (PRMs) excel at providing step-by-step verification for large language model (LLM) outputs in domains like mathematics and coding. However, their a…