3 papers
cs.HC2025
MobileViews: A Million-scale and Diverse Mobile GUI Dataset
Longxi Gao, Li Zhang, Shihe Wang +6
Visual language models (VLMs) empower mobile GUI agents to interpret complex mobile screens and respond to user requests. Training such capable agents requires large-scale, high-qu…
cs.AI2025
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents
Yunhe Yan, Shihe Wang, Jiajun Du +12
(M)LLM-powered computer use agents (CUA) are emerging as a transformative technique to automate human-computer interaction. However, existing CUA benchmarks predominantly target GU…
cs.AI2024
DroidCall: A Dataset for LLM-powered Android Intent Invocation
Weikai Xie, Li Zhang, Shihe Wang +2
The growing capabilities of large language models in natural language understanding significantly strengthen existing agentic systems. To power performant on-device mobile agents f…