1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Xiong Jun Wu, Zhenduo Zhang, ZuJie Wen +11
Training large reasoning models (LRMs) with reinforcement learning in STEM domains is hindered by the scarcity of high-quality, diverse, and verifiable problem sets. Existing synth…