1 paper
Kun Yang, Zikang chen, Yanmeng Wang +4
As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved reasoning performance. However, exi…