1 paper
Woojeong Kim, Ziyi Yang, Jing Nathan Yan +1
Reinforcement learning (RL) is the dominant paradigm for post-training large language models. However, in the online, on-policy setting, rollout generation dominates the computatio…