1 paper
Dingwei Chen, Zefang Zong, Zhipeng Ma +5
Reinforcement learning for agentic large language models (LLMs) typically relies on a sparse, trajectory-level outcome reward, making it difficult to evaluate the contribution of i…