1 paper · 1 filter
Yawen Shao, Jie Xiao, Kai Zhu +4
Group Relative Policy Optimization (GRPO) has proven highly effective in enhancing the alignment capabilities of Large Language Models (LLMs). However, current adaptations of GRPO…