1 paper · 1 filter
Can Xie, Yuyi Zhou, Wen Yang +5
Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the reusable exploration patterns embedded in i…