From the 1 of 8 linked papers with an AI index.
8 papers
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
Hanzhang Zhou, Panrong Tong, Xu Zhang +13
The paper introduces Qwen-UI-Agent, a foundation model for GUI agents that can operate across mobile, desktop, web, and search environments, combining GUI actions with CLI commands…
What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States
Chen Liu, Ling Chen, Hanzhang Zhou +7
Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications. Existing methods treat memory l…
One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Grounding
Chen Liu, Ling Chen, Hanzhang Zhou +5
MLLM-based GUI grounding methods commonly formulate target localization as autoregressive coordinate generation, enabling models to leverage the strong instruction-following and se…
WebGen-R1: Incentivizing Large Language Models to Generate Functional and Aesthetic Websites with Reinforcement Learning
Juyong Jiang, Chenglin Cai, Chansung Park +4
While Large Language Models (LLMs) excel at function-level code generation, project-level tasks such as generating functional and visually aesthetic multi-page websites remain high…
FedGUI: Benchmarking Federated GUI Agents across Heterogeneous Platforms, Devices, and Operating Systems
Wenhao Wang, Haoting Shi, Mengying Yuan +7
Training GUI agents with traditional centralized methods faces significant cost and scalability challenges. Federated learning (FL) offers a promising solution, yet its potential i…
MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented Environments
Quyu Kong, Xu Zhang, Zhenyu Yang +10
Among existing online mobile-use benchmarks, AndroidWorld has emerged as the dominant benchmark due to its reproducible environment and deterministic evaluation; however, recent ag…