1 paper · 1 filter
Keyu Duan, Zichen Liu, Xin Mao +5
Process Reward Models (PRMs) provide step-level supervision to large language models (LLMs), but scaling up training data annotation remains challenging for both humans and LLMs. T…