1 paper
Shuo Yang, Soyeon Caren Han, Xueqi Ma +3
LLM-based agents depend on effective tool-use policies to solve complex tasks, yet optimizing these policies remains challenging due to delayed supervision and the difficulty of cr…