1 paper · 1 filter
Ranxu Zhang, Guinan Chen, Chenshaodong +5
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions…