3 papers
cs.AI2026
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
Shaokang Wang, Pei Fu, Ruoceng Zhang +7
While Large Vision-Language Models (LVLMs) have significantly advanced GUI agents' capabilities in parsing textual instructions, interpreting screen content, and executing tasks, a…
cs.CV2026
Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
Shaojie Zhang, Pei Fu, Ruoceng Zhang +8
Autonomous graphical user interface (GUI) agents rely on accurate GUI grounding, which maps language instructions to on-screen coordinates, to execute user commands. However, curre…
cs.LG2026
Beyond Binary: Reframing GUI Critique as Continuous Semantic Alignment
Yuchen Sun, Pei Fu, Shaojie Zhang +6
Test-Time Scaling (TTS), which samples multiple candidate actions and ranks them via a Critic Model, has emerged as a promising paradigm for generalist GUI agents. Its efficacy thu…