1 paper
Prasanth YSS, Zhichen Ren, Rasa Hosseinzadeh +6
Reinforcement learning with verifiable rewards (RLVR) improves language-model reasoning, but GRPO-style optimization remains prone to collapse. We analyse this instability through…