2 papers
cs.CV2026
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
Shrinidhi Kumbhar, Haofu Liao, Srikar Appalaraju +1
Autoregressive (AR) vision-language models (VLMs) have long dominated multimodal understanding, reasoning, and graphical user interface (GUI) grounding. Recently, discrete diffusio…
cs.CV2025
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
Joonhyung Park, Peng Tang, Sagnik Das +4
Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by advances in large Vision Language…