10 papers
UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
Yuxiang Chai, Han Xiao, Xinyu Fu +3
Recent advances in mobile GUI agents have shown strong potential for automating mobile tasks, but most effective systems still depend on large vision-language models for screenshot…
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
Xudong Lu, Xueying Li, Annan Wang +8
We introduce OmniInteract, a streaming benchmark for real-time omnimodal large language models evaluated through native online inference over audio-visual streams. Unlike offline v…
Rethinking VLM Representation for VLA Initialization
Weifeng Lin, Siyuan Huang, Hao Li +5
Vision-Language-Action (VLA) models widely adopt pretrained Vision-Language Models (VLMs) as policy backbones, yet it remains unclear what kind of pretrained VLM representation is…
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
Dian Zheng, Manyuan Zhang, Hongyu Li +4
Currently, enhancing Unified Multimodal Models (UMMs) with image understanding, generation, and editing capabilities mainly relies on mixed multi-task training. Due to inherent tas…
PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation Agents
Yuxiang Chai, Shunye Tang, Han Xiao +2
Current Graphical User Interface (GUI) agents operate primarily under a reactive paradigm: a user must provide an explicit instruction for the agent to execute a task. However, an…
UI-Mem: Self-Evolving Experience Memory for Online Reinforcement Learning in Mobile GUI Agents
Han Xiao, Guozhi Wang, Hao Wang +7
Online Reinforcement Learning (RL) offers a promising paradigm for enhancing GUI agents through direct environment interaction. However, its effectiveness is severely hindered by i…