2 papers
cs.CL2025
More Vulnerable than You Think: On the Stability of Tool-Integrated LLM Agents
Weimin Xiong, Ke Wang, Yifan Song +4
Current evaluations of tool-integrated LLM agents typically focus on end-to-end tool-usage evaluation while neglecting their stability. This limits their real-world applicability,…
cs.CL2024
AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories
Yifan Song, Weimin Xiong, Xiutian Zhao +6
Fine-tuning on agent-environment interaction trajectory data holds significant promise for surfacing generalized agent capabilities in open-source large language models (LLMs). In…