1 paper
Minseo Kim, Minjae Lee, Seunghyuk Oh +7
Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling strong reasoning and agentic capabilities. However, rollout generation remains a d…