activity
20242026
collaborators

7 papers

cs.CV2026

Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding

Shrinidhi Kumbhar, Haofu Liao, Srikar Appalaraju +1

Autoregressive (AR) vision-language models (VLMs) have long dominated multimodal understanding, reasoning, and graphical user interface (GUI) grounding. Recently, discrete diffusio…

cs.CL2025

Turbocharging Web Automation: The Impact of Compressed History States

Xiyue Zhu, Peng Tang, Haofu Liao +1

Language models have led to a leap forward in web automation. The current web automation approaches take the current web state, history actions, and language instruction as inputs…

cs.CV2025

R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding

Joonhyung Park, Peng Tang, Sagnik Das +4

Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by advances in large Vision Language…

cs.CV2025

VisFocus: Prompt-Guided Vision Encoders for OCR-Free Dense Document Understanding

Ofir Abramovich, Niv Nayman, Sharon Fogel +7

In recent years, notable advancements have been made in the domain of visual document understanding, with the prevailing architecture comprising a cascade of vision and language mo…

cs.CV2025

Enhancing Vision-Language Pre-training with Rich Supervisions

Yuan Gao, Kunyu Shi, Pengkai Zhu +7

We propose Strongly Supervised pre-training with ScreenShots (S4) - a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. U…

cs.CV2024

DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models

Sungnyun Kim, Haofu Liao, Srikar Appalaraju +6

Visual document understanding (VDU) is a challenging task that involves understanding documents across various modalities (text and image) and layouts (forms, tables, etc.). This s…