1 paper · 1 filter
Liming Pu, Xiaoxia Li, Yifu Liu +2
Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-…