2 papers
cs.AI2025
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
Yunhe Yan, Shihe Wang, Jiajun Du +12
(M)LLM-powered computer use agents (CUA) are emerging as a transformative technique to automate human-computer interaction. However, existing CUA benchmarks predominantly target GU…
cs.HC2024
LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task Automation
Li Zhang, Shihe Wang, Xianqing Jia +5
The emergent large language/multimodal models facilitate the evolution of mobile agents, especially in mobile UI task automation. However, existing evaluation approaches, which rel…