1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Tianle Wang, Zhaoyang Wang, Guangchen Lan +4
Reinforcement learning (RL) has been applied to improve large language model (LLM) reasoning, yet the systematic study of how training scales with task difficulty has been hampered…