3 citations · 17 across the 14 of their papers we have counts for
1 paper · 1 filter
Peng Yu, Zeyuan Zhao, Shao Zhang +3
Large language models (LLMs) have achieved significant advancements in reasoning capabilities through reinforcement learning (RL) via environmental exploration. As the intrinsic pr…