Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Towards Robust Process Reward Modeling via Noise-aware Learning
Bin Xie, Bingbing Xu, Xueyun Tian +2
Process Reward Models (PRMs) have achieved strong results in complex reasoning, but are bottlenecked by costly process-level supervision. A widely used alternative, Monte Carlo Est…
cs.CL2025
From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment
Bin Xie, Bingbing Xu, Yige Yuan +2
Inference-time alignment methods have gained significant attention for their efficiency and effectiveness in aligning large language models (LLMs) with human preferences. However,…