4 papers
Multimodal Graph RAG for Long-range Visually Rich Document Understanding
Yi-Cheng Wang, Chu-Song Chen
Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue by the limited context window.…
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
Chi-Hsiang Hsiao, Yi-Cheng Wang, Tzung-Sheng Lin +2
Retrieval-augmented generation (RAG) enables large language models (LLMs) to dynamically access external information, which is powerful for answering questions over previously unse…
Document-Level Numerical Reasoning across Single and Multiple Tables in Financial Reports
Yi-Cheng Wang, Wei-An Wang, Chu-Song Chen
Despite the strong language understanding abilities of large language models (LLMs), they still struggle with reliable question answering (QA) over long, structured documents, part…
Align-GRAG: Anchor and Rationale Guided Dual Alignment for Graph Retrieval-Augmented Generation
Derong Xu, Pengyue Jia, Xiaopeng Li +9
Despite the strong abilities, large language models (LLMs) still suffer from hallucinations and reliance on outdated knowledge, raising concerns in knowledge-intensive tasks. Graph…