agent evaluation 1asynchronous runtime 1benchmarking 1large language models 1software infrastructure 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.AI2026
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
Kai Chen, Zichen Ding, Jiaye Ge +22
The paper presents AgentCompass, an open‑source infrastructure that standardizes and simplifies the evaluation of large‑language‑model based autonomous agents by separating benchma…
cs.LG2026
Looped World Models
Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28
Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deeper models are expensive to deploy and prone to compounding error…