1 paper · 1 filter
Mohamed Sana, Nicola Piovesan, Antonio De Domenico +2
We investigate a narrow but common failure mode of GRPO-style reinforcement learning in the context of sparse verifiable rewards: early updates contain more responses with negative…