4 papers
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Prince Zizhuang Wang, Aojie Yuan, Haiyue Zhang +3
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on t…
FORTIS: Benchmarking Over-Privilege in Agent Skills
Shawn Li, Chenxiao Yu, Han Wang +8
Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task execution. This layer is widely treated as…
The Sim-to-Real Gap of Foundation Model Agents: A Unified MDP Perspective
Xiaoou Liu, Tiejin Chen, Weibo Li +2
Foundation model agents are increasingly deployed for real-world decision-making, but suffer from the sim-to-real gap. While robotics and classical control have mature frameworks t…
Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
Xiaolin Zhou, Zheng Luo, Yicheng Gao +4
Recent advances in Large Language Models (LLMs) have incentivized the development of LLM-as-a-judge, an application of LLMs where they are used as judges to decide the quality of a…