1 paper · 1 filter
Boyuan Meng, Peihua Bao, Hong Liu +4
Agentic reinforcement learning (RL) often produces irregular rollout trees with shared histories. Training root-to-leaf trajectories independently recomputes these shared prefixes.…