1 paper · 1 filter
Duc Anh Le, Tien-Phat Nguyen, Thien Huu Nguyen +2
Reinforcement learning with verifiable rewards helps reasoning, but GRPO-style methods stall on hard prompts where all sampled rollouts fail. SORT adds a repair update for those fa…