1 paper · 1 filter
Wenke Huang, Quan Zhang, Yiyang Fang +11
Recent advances in reinforcement learning for foundation models, such as Group Relative Policy Optimization (GRPO), have significantly improved the performance of foundation models…