1 citations · 1 across the 8 of their papers we have counts for
8 papers
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang +3
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on t…
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
Zheng Luo, T Pranav Kutralingam, Ogochukwu N Okoani +3
Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent work reports strong tool-calling perform…
FORTIS: Benchmarking Over-Privilege in Agent Skills
Shawn Li, Chenxiao Yu, Han Wang +8
Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This layer is widely treated as…
Language Shapes Mental Health Evaluations in Large Language Models
Jiayi Xu, Xiyang Hu
Multilingual large language models (LLMs) are increasingly used in socially sensitive mental health contexts, including support chatbots, screening, and content moderation. This ra…
The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective
Xiaoou Liu, Tiejin Chen, Weibo Li +2
Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical control have mature frameworks t…
Counterfactual Trace Auditing of LLM Agent Skills
Xiaolin Zhou, Jinbo Liu, Li Li +2
Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchmarks report only pass rate befor…