Evaluation and Benchmarking of LLM Agents: A Survey
arXiv:2507.21504 · doi:10.1145/3711896.3736570
Abstract
The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation, introducing a two-dimensional taxonomy that organizes existing work along (1) evaluation objectives -- what to evaluate, such as agent behavior, capabilities, reliability, and safety -- and (2) evaluation process -- how to evaluate, including interaction modes, datasets and benchmarks, metric computation methods, and tooling. In addition to taxonomy, we highlight enterprise-specific challenges, such as role-based access to data, the need for reliability guarantees, dynamic and long-horizon interactions, and compliance, which are often overlooked in current research. We also identify future research directions, including holistic, more realistic, and scalable evaluation. This work aims to bring clarity to the fragmented landscape of agent evaluation and provide a framework for systematic assessment, enabling researchers and practitioners to evaluate LLM agents for real-world deployment.
References in corpus (12)
- Gorilla: Large Language Model Connected with Massive APIs
- Understanding the planning of LLM agents: A survey
- Magentic-One: A Generalist Multi-Agent System for Solving Complex Tasks
- A Survey on the Memory Mechanism of Large Language Model based Agents
- -bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
- WebCanvas: Benchmarking Web Agents in Online Environments
- AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents
- AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
- FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
- IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems