6 papers
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
Weiming Li, Yan Shao, Jing Yang +6
Graphical user interface (GUI) grounding is a fundamental task for building GUI agents. However, general vision-language models (VLMs) struggle with this task due to a lack of spec…
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
Jinhao Jing, Zheng Ma, Jinwei Liang +9
Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address thi…
Architecture Matters More Than Scale: A Comparative Study of Retrieval and Memory Augmentation for Financial QA Under SME Compute Constraints
Jianan Liu, Jing Yang, Xianyou Li +4
The rapid adoption of artificial intelligence (AI) and large language models (LLMs) is transforming financial analytics by enabling natural language interfaces for reporting, decis…
TA-Mem: Tool-Augmented Autonomous Memory Retrieval for LLM in Long-Term Conversational QA
Mengwei Yuan, Jianan Liu, Jing Yang +4
Large Language Model (LLM) has exhibited strong reasoning ability in text-based contexts across various domains, yet the limitation of context window poses challenges for the model…
DynaRAG: Bridging Static and Dynamic Knowledge in Retrieval-Augmented Generation
Penghao Liang, Mengwei Yuan, Jianan Liu +4
We present DynaRAG, a retrieval-augmented generation (RAG) framework designed to handle both static and time-sensitive information needs through dynamic knowledge integration. Unli…
DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts
Yujing Lu, Ling Zhong, Jing Yang +5
Chart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test su…