3 papers
cs.AI2026
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Zedong Yu, Qianxing Li, Zhi Gao +8
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as cl…
cs.AI2026
GUI Knowledge Bench: Revealing the Knowledge Gap of VLMs in GUI Tasks
Chenrui Shi, Zedong Yu, Zhi Gao +7
Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledg…
cs.LG2025
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
Pengxiang Li, Zechen Hu, Zirui Shang +15
Vision-language model (VLM) based GUI agents show promise for automating complex desktop and mobile tasks, but face significant challenges in applying reinforcement learning (RL):…