1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Jiecheng Zhou, Qinghao Hu, Peng Sun +2
Large language model (LLM) agents are increasingly trained with reinforcement learning in long-horizon, sandboxed environments. Unlike conventional RL, agentic RL couples GPU-inten…