most citedA Survey on (M)LLM-Based GUI Agents

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CV2025

SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation

Siqi Chen, Xinyu Dong, Haolei Xu +10

Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of comp…

cs.CV2025

ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models

Dingming Li, Hongxing Li, Zixuan Wang +9

Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring c…

cs.HC20251 cited

A Survey on (M)LLM-Based GUI Agents

Fei Tang, Haolei Xu, Hang Zhang +12

Graphical User Interface (GUI) Agents have emerged as a transformative paradigm in human-computer interaction, evolving from rule-based automation scripts to sophisticated AI-drive…

cs.AI2025

Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems

Fei Tang, Yongliang Shen, Hang Zhang +7

Humans can flexibly switch between different modes of thinking based on task complexity: from rapid intuitive judgments to in-depth analytical understanding. However, current Graph…

cs.CL2025

Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Wenqi Zhang, Mengna Wang, Gangao Liu +10

Recent advances in deep thinking models have demonstrated remarkable reasoning capabilities on mathematical and coding tasks. However, their effectiveness in embodied domains which…