1 citations · 1 across the 4 of their papers we have counts for
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Kean Shi, Zihang Li, Tianyi Ma +13
Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex environments, such as web browse…
cs.AI2026
EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
Haiyang Shen, Xuanzhong Chen, Wendong Xu +3
Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assessment. This leaves out a basi…