1 paper
Ran Li, Shimin Di, Yuchen Liu +3
Previous study suggest that powerful Large Language Models (LLMs) trained with Reinforcement Learning with Verifiable Rewards (RLVR) only refines reasoning path without improving t…