1 paper
Yinghui Chi, Yuanhong Wang
Process reward models (PRMs) require supervision that identifies not only whether a reasoning trajectory is correct, but also where the reasoning process first becomes unsupported…