3 papers
cs.CV2026
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
Shrinidhi Kumbhar, Haofu Liao, Srikar Appalaraju +1
Autoregressive (AR) vision-language models (VLMs) have long dominated multimodal understanding, reasoning, and graphical user interface (GUI) grounding. Recently, discrete diffusio…
cs.CL2025
Turbocharging Web Automation: The Impact of Compressed History States
Xiyue Zhu, Peng Tang, Haofu Liao +1
Language models have led to a leap forward in web automation. The current web automation approaches take the current web state, history actions, and language instruction as inputs…
cs.CV2024
DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models
Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6
Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…