3 papers
cs.HC2026
Benchmarking LLM Tool-Use in the Wild
Peijie Yu, Wei Liu, Yifan Yang +4
Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently wild, being intricate,…
cs.AI2025
-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
Peijie Yu, Yifan Yang, Jinjian Li +4
Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely…
cs.AI2025
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
Peijie Yu, Yifan Yang, Jinjian Li +4
Large language models (LLMs) demonstrate strong potential as agents for tool invocation due to their advanced comprehension and planning capabilities. Users increasingly rely on LL…