5 papers
Toward Scalable Terminal Task Synthesis via Skill Graphs
Zhiyuan Fan, Tinghao Yu, Yuanjun Cai +8
Terminal agents have demonstrated strong potential for autonomous command-line execution, yet their training remains constrained by the scarcity of high-quality and diverse executi…
ClawLess: A Security Model of AI Agents
Hongyi Lu, Nian Liu, Shuai Wang +1
Autonomous AI agents powered by Large Language Models can reason, plan, and execute complex tasks, but their ability to autonomously retrieve information and run code introduces si…
Benchmarking LLM Tool-Use in the Wild
Peijie Yu, Wei Liu, Yifan Yang +4
Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inherently wild, being intricate,…
-Bench: The Things Real Disturbing LLM based Agent in Multi-Tasking
Peijie Yu, Yifan Yang, Jinjian Li +4
Agents based on large language models leverage tools to modify environments, revolutionizing how AI interacts with the physical world. Unlike traditional NLP tasks that rely solely…
Multi-Mission Tool Bench: Assessing the Robustness of LLM based Agents through Related and Dynamic Missions
Peijie Yu, Yifan Yang, Jinjian Li +4
Large language models (LLMs) demonstrate strong potential as agents for tool invocation due to their advanced comprehension and planning capabilities. Users increasingly rely on LL…