Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Junzhi Chen, Harsh Trivedi, Jane Pan +4
Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification que…
cs.AI2025
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
Sayash Kapoor, Benedikt Stroebl, Peter Kirgis +28
AI agents have been developed for complex real-world tasks from coding to customer service. But AI agent evaluations suffer from many challenges that undermine our understanding of…