1 paper
Yi Yang, Cong Qin, Xiaodan Liu +8
Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens…