1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Chen Wang, Hexuan Deng, Yining Zhang +5
Reinforcement learning with verifiable rewards improves LLM reasoning but often induces overthinking, where models generate unnecessarily long reasoning traces. Existing methods ma…