activity
20242026
most citedA Survey on (M)LLM-Based GUI Agents

1 citations · 1 across the 13 of their papers we have counts for

collaborators
Showing 2026Show all

8 papers · 1 filter

cs.LG2026

Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

Bofan Chen, Boxuan Zhang, Fei Tang +7

GUI agents execute long-horizon tasks on dynamic graphical user interfaces, where pop-ups, delayed loads, and relocated widgets routinely invalidate plans fixed before execution. R…

cs.CV2026

Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding

Yizhou Liu, Fei Tang, Yuchen Yan +8

Graphical User Interface (GUI) grounding is essential for autonomous agents to map natural language instructions to precise screen coordinates. However, existing supervised fine-tu…

cs.CL2026

BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes

Fei Tang, Huawen Shen, Zhiqiong Lu +7

Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high…

cs.CL2026

PhoneWorld: Scaling Phone-Use Agent Environments

Yuxuan Liu, Xin Lai, Junyi Li +21

A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Existing mobile-agent benchmarks…

cs.CV2026

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding

Fei Tang, Bofan Chen, Zhengxi Lu +8

GUI grounding, which localizes interface elements from screenshots given natural language queries, remains challenging for small icons and dense layouts. Test-time zoom-in methods…

cs.LG2026

UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization

Zhengxi Lu, Fei Tang, Guangyi Liu +8

MLLM-based GUI agents have demonstrated strong capabilities in complex user interface interaction tasks. However, long-horizon scenarios remain challenging, as these agents are bur…