most citedA Survey on (M)LLM-Based GUI Agents

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CV2025

SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation

Siqi Chen, Xinyu Dong, Haolei Xu +10

Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of comp…

cs.CV2025

ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models

Dingming Li, Hongxing Li, Zixuan Wang +9

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring c…

cs.CV2025

Generalized Visual Relation Detection with Diffusion Models

Kaifeng Gao, Siqi Chen, Hanwang Zhang +3

Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance,…

cs.HC20251 cited

A Survey on (M)LLM-Based GUI Agents

Fei Tang, Haolei Xu, Hang Zhang +12

Graphical User Interface (GUI) Agents have emerged as a transformative paradigm in human-computer interaction, evolving from rule-based automation scripts to sophisticated AI-drive…

cs.AI2025

Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems

Fei Tang, Yongliang Shen, Hang Zhang +7

Humans can flexibly switch between different modes of thinking based on task complexity: from rapid intuitive judgments to in-depth analytical understanding. However, current Graph…