4 citations · 4 across the 1 of their papers we have counts for
1 paper · 1 filter
Guanting Dong, Yifei Chen, Xiaoxi Li +7
Recently, large language models (LLMs) have shown remarkable reasoning capabilities via large-scale reinforcement learning (RL). However, leveraging the RL algorithm to empower eff…