4 papers
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
Weiming Li, Yan Shao, Jing Yang +6
Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with this task due to a lack of spec…
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
Yujing Lu, Ling Zhong, Jing Yang +5
Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test su…
Towards Building a Robust Knowledge Intensive Question Answering Model with Large Language Models
Xingyun Hong, Yan Shao, Zhilin Wang +2
The development of LLMs has greatly enhanced the intelligence and fluency of question answering, while the emergence of retrieval enhancement has enabled models to better utilize e…
Large Language Models Understand Layout
Weiming Li, Manni Duan, Dong An +1
Large language models (LLMs) demonstrate extraordinary abilities in a wide range of natural language processing (NLP) tasks. In this paper, we show that, beyond text understanding…