AgentArcEval: An Architecture Evaluation Method for Foundation Model based Agents
arXiv:2510.21031 · doi:10.1016/j.jss.2025.112656
Abstract
The emergence of foundation models (FMs) has enabled the development of highly capable and autonomous agents, unlocking new application opportunities across a wide range of domains. Evaluating the architecture of agents is particularly important as the architectural decisions significantly impact the quality attributes of agents given their unique characteristics, including compound architecture, autonomous and non-deterministic behaviour, and continuous evolution. However, these traditional methods fall short in addressing the evaluation needs of agent architecture due to the unique characteristics of these agents. Therefore, in this paper, we present AgentArcEval, a novel agent architecture evaluation method designed specially to address the complexities of FM-based agent architecture and its evaluation. Moreover, we present a catalogue of agent-specific general scenarios, which serves as a guide for generating concrete scenarios to design and evaluate the agent architecture. We demonstrate the usefulness of AgentArcEval and the catalogue through a case study on the architecture evaluation of a real-world tax copilot, named Luna.
References in corpus (14)
- On the Opportunities and Risks of Foundation Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Self-Refine: Iterative Refinement with Self-Feedback
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Cognitive Architectures for Language Agents
- MemGPT: Towards LLMs as Operating Systems
- Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents
- Building Cooperative Embodied Agents Modularly with Large Language Models
- SurrealDriver: Designing LLM-powered Generative Driver Agent Framework based on Human Drivers' Driving-thinking Data
- OpenAgents: An Open Platform for Language Agents in the Wild
- InterAct: Exploring the Potentials of ChatGPT as a Cooperative Agent
- TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems