1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Haipeng Luo, Qingfeng Sun, Songli Wu +4
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from poli…