2 papers
cs.AI2026
Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios
Zuoyu Zhang, Yancheng Zhu
Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environments. Yet existing safety be…
cs.HC2026
Enhancing Tool Calling in LLMs with the International Tool Calling Dataset
Zuoyu Zhang, Yancheng Zhu
Tool calling allows large language models (LLMs) to interact with external systems like APIs, enabling applications in customer support, data analysis, and dynamic content generati…