2 papers
cs.CL2026
Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
Ziheng Li, Liu Kang, Feng Xiao +7
Group Relative Policy Optimization (GRPO) has emerged as a promising critic-free reinforcement learning paradigm for reasoning tasks. However, standard GRPO employs a coarse-graine…
cs.LG2026
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
Haonan Song, Qingchen Xie, Huan Zhu +12
Generative Reward Models (GRMs) have demonstrated strong performance in reward modeling, due to their interpretability and potential for refinement through reinforcement learning (…