1 paper
Yingxuan Zhuang, Binhe Yu, Jingxiao Yang +6
Reinforcement learning for LLM agents involves two distinct optimization di- mensions: how environment feedback is exploited within a trajectory, and how complete trajectories are…