2 papers
cs.CL2025
Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
Chen Zheng, Yiyuan Ma, Yuan Yang +11
The development of alignment and reasoning capabilities in large language models has seen remarkable progress through two paradigms: instruction tuning and reinforcement learning f…
cs.LG2025
Prefix Grouper: Efficient GRPO Training through Shared-Prefix Forward
Zikang Liu, Tongtian Yue, Yepeng Tang +5
Group Relative Policy Optimization (GRPO) enhances policy learning by computing gradients from relative comparisons among candidate outputs that share a common input prefix. Despit…