agent evaluation 1asynchronous runtime 1benchmarking 1large language models 1software infrastructure 1
From the 1 of 11 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Kai Chen, Zichen Ding, Jiaye Ge +22
The paper presents AgentCompass, an open‑source infrastructure that standardizes and simplifies the evaluation of large‑language‑model based autonomous agents by separating benchma…
cs.AI2025
Are Your LLMs Capable of Stable Reasoning?
Junnan Liu, Hongwei Liu, Linchen Xiao +6
The rapid advancement of large language models (LLMs) has shown remarkable progress in complex reasoning tasks. However, a significant disparity exists between benchmark performanc…