12 papers
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
Chenrui Shi, Yuwei Wu, Yang Liu +5
The paper introduces an Interactive Reward Agent that evaluates GUI task completion by proposing conditions and verifying them using system, application, and GUI tools, and demonst…
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Zedong Yu, Qianxing Li, Zhi Gao +8
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user instructions through GUI actions such as cl…
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
Rui Xie, Zhi Gao, Chenrui Shi +3
Large vision-language models have endowed GUI agents with strong general capabilities for interface understanding and interaction. However, due to insufficient exposure to domain-s…
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
Pengxiang Li, Zhi Gao, Bofei Zhang +8
Multimodal agents, which integrate a controller e.g., a vision language model) with external tools, have demonstrated remarkable capabilities in tackling complex multimodal tasks.…
Benchmarking and Improving GUI Agents in High-Dynamic Environments
Enqi Liu, Liyuan Pan, Zhi Gao +5
Recent advancements in Graphical User Interface (GUI) agents have predominantly focused on training paradigms like supervised fine-tuning (SFT) and reinforcement learning (RL). How…
When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations
Kailin Jiang, Yuntao Du, Yukai Ding +7
Large Multimodal Models (LMMs) store vast amounts of pretrained knowledge but struggle to remain aligned with real-world updates, making it difficult to avoid capability degradatio…