3 papers
cs.CL2026
GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent Environments
Lin Fu, Zheyuan Yang, Tianhui Zhang +5
GUI world models are increasingly evaluated as one-step next-screen predictors, yet their intended use is often as multi-step environments for GUI agents. This mismatch leaves a ke…
cs.AI2026
Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
Guo Gan, Yilun Zhao, Cong Chen +5
GUI agents often encounter dynamic anomalies when deployed on Android devices, from unexpected pop-ups to action misuse, yet existing benchmarks lack systematic evaluation of agent…
cs.LG2026
Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple Actions
Guo Gan, Yuxuan Ding, Cong Chen +3
Online reinforcement learning (RL) serves as an effective method for enhancing the capabilities of Android agents. However, guiding agents to learn through online interaction is pr…