1 citations · 1 across the 3 of their papers we have counts for
5 papers
SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
Siqi Chen, Xinyu Dong, Haolei Xu +10
Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of comp…
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dingming Li, Hongxing Li, Zixuan Wang +9
Vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and reasoning about visual content, but significant challenges persist in tasks requiring c…
Generalized Visual Relation Detection with Diffusion Models
Kaifeng Gao, Siqi Chen, Hanwang Zhang +3
Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance,…
A Survey on (M)LLM-Based GUI Agents
Fei Tang, Haolei Xu, Hang Zhang +12
Graphical User Interface (GUI) Agents have emerged as a transformative paradigm in human-computer interaction, evolving from rule-based automation scripts to sophisticated AI-drive…
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
Fei Tang, Yongliang Shen, Hang Zhang +7
Humans can flexibly switch between different modes of thinking based on task complexity: from rapid intuitive judgments to in-depth analytical understanding. However, current Graph…