2 papers
cs.CV2026
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
Weiming Li, Yan Shao, Jing Yang +6
Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with this task due to a lack of spec…
cs.CL2026
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
Yujing Lu, Ling Zhong, Jing Yang +5
Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test su…