collaborators

7 papers

cs.CL2026

SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluations

Shuaiqi Wang, Aadyaa Maddi, Zinan Lin +1

Today, tool-calling agents are commonly evaluated or tested on static datasets of execution traces, including input commands, agent responses, and associated tool calls. However, i…

cs.SE2026

Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode

Zimo Ji, Zongjie Li, Wenyuan Jiang +2

Claude Code's auto mode is the first deployed permission system for AI coding agents, using a two-stage transcript classifier to gate dangerous tool calls. Anthropic reports a 0.4%…

cs.CR2026

ClawLess: A Security Model of AI Agents

Hongyi Lu, Nian Liu, Shuai Wang +1

Autonomous AI agents powered by Large Language Models can reason, plan, and execute complex tasks, but their ability to autonomously retrieve information and run code introduces si…

cs.CL2025

ACEBench: Who Wins the Match Point in Tool Usage?

Chen Chen, Xinlong Hao, Weiwen Liu +13

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex…

cs.LG2025

ToolACE: Winning the Points of LLM Function Calling

Weiwen Liu, Xu Huang, Xingshan Zeng +24

Function calling significantly extends the application boundary of large language models, where high-quality and diverse training data is critical for unlocking this capability. Ho…

cs.AI2025

SPA-Bench: A Comprehensive Benchmark for SmartPhone Agent Evaluation

Jingxuan Chen, Derek Yuen, Bin Xie +14

Smartphone agents are increasingly important for helping users control devices efficiently, with (Multimodal) Large Language Model (MLLM)-based approaches emerging as key contender…