1 paper
Feifan Song, Shaohang Wei, Bofei Gao +8
Large reasoning models (LRMs) boosted by Reinforcement Learning from Verifier Reward (RLVR) have shown great power in problem solving, yet they often cause overthinking: excessive,…