1 paper
Hao Ye, Jisheng Dang, Junfeng Fang +7
Recent extensive research has demonstrated that the enhanced reasoning capabilities acquired by models through Reinforcement Learning with Verifiable Rewards (RLVR) are primarily c…