1 paper
Xingyu Shen, Huishuai Zhang, Peng Li +2
Reinforcement Learning with Verifiable Rewards (RLVR) significantly improves LLM reasoning but often causes a drop in policy entropy, leading to narrowed reasoning coverage and deg…