3 papers
cs.CL2025
-GRPO: Unifying the GRPO Frameworks with Learnable Token Preferences
Yining Wang, Jinman Zhao, Chuangxin Zhao +3
Reinforcement Learning with Human Feedback (RLHF) has been the dominant approach for improving the reasoning capabilities of Large Language Models (LLMs). Recently, Reinforcement L…
cs.LG2025
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
Yuhui Sun, Xiyao Wang, Zixi Li +6
Recent alignment methods based on Direct Preference Optimization (DPO) reformulate preference learning as supervised optimization over pairwise comparisons, offering improved effic…
cs.AI2025
Coarse-to-Fine Process Reward Modeling for Mathematical Reasoning
Yulan Hu, Sheng Ouyang, Jinman Zhao +1
The Process Reward Model (PRM) plays a crucial role in mathematical reasoning tasks, requiring high-quality supervised process data. However, we observe that reasoning steps genera…