8 citations · 22 across the 13 of their papers we have counts for
1 paper · 1 filter
Wenke Huang, Quan Zhang, Yiyang Fang +11
Recent advances in reinforcement learning for foundation models, such as Group Relative Policy Optimization (GRPO), have significantly improved the performance of foundation models…