1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Xiao Liang, Zhong-Zhi Li, Yeyun Gong +5
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for training large language models (LLMs) on complex reasoning tasks, such as mathematical problem solvin…