19 citations · 19 across the 5 of their papers we have counts for
1 paper · 1 filter
Zijian Wu, Lingkai Kong, Wenwei Zhang +12
Large language models (LLMs) have achieved significant progress in solving complex reasoning tasks by Reinforcement Learning with Verifiable Rewards (RLVR). This advancement is als…