1 paper
Mianjie Yu, Zizhao Mo, Huanyu Qu +8
In modern agentic AI training, the system bottleneck is shifting from rollout to update. Emerging sampling strategies such as tree-structured and step-wise RL greatly increase trai…