1 paper
Weimin Xiong, Ke Wang, Yifan Song +4
Current evaluations of tool-integrated LLM agents typically focus on end-to-end tool-usage evaluation while neglecting their stability. This limits their real-world applicability,…