3 papers
cs.CV2025
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
Weiming Li, Yan Shao, Jing Yang +6
Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with this task due to a lack of spec…
cs.CL2025
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
Yujing Lu, Ling Zhong, Jing Yang +5
Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test su…
cs.CL2024
Large Language Models Understand Layout
Weiming Li, Manni Duan, Dong An +1
Large language models (LLMs) demonstrate extraordinary abilities in a wide range of natural language processing (NLP) tasks. In this paper, we show that, beyond text understanding…