2 papers
cs.LG2025★ 1 cited
ScalePRM: Training Process Reward Models by Scaling Verification Compute Without Ground Truth
Salman Rahman, Sruthi Gorantla, Arpit Gupta +3
Training process reward models (PRMs) requires step-level correctness labels, obtained either through expensive human annotation or by relying on ground-truth answers, limiting the…
cs.LG2025
SABER: Small Actions, Big Errors -- Safeguarding Mutating Steps in LLM Agents
Alejandro Cuadron, Pengfei Yu, Yang Liu +1
Despite rapid progress in LLM agents, performance on long-horizon, tool-using tasks remains fragile. To better understand this fragility, we ask a simple question: \emph{do all act…