2 papers
cs.LG2026
BAGEN: Are LLM Agents Budget-Aware?
Yuxiang Lin, Zihan Wang, Mengyang Liu +9
While agents are increasingly spending more resources, today agent cost is mostly measured only after execution. A Budget-Aware Agent (BAGEN) should treat budget as an active contr…
cs.AI2025
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
Yunhe Yan, Shihe Wang, Jiajun Du +12
(M)LLM-powered computer use agents (CUA) are emerging as a transformative technique to automate human-computer interaction. However, existing CUA benchmarks predominantly target GU…