1 paper
Yongyu Mu, Jiali Zeng, Fandong Meng +2
Through encouraging self-exploration, reinforcement learning from verifiable rewards (RLVR) has significantly advanced the mathematical reasoning capabilities of large language mod…