4 papers
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
Yihao Ding, Siwen Luo, Yue Dai +6
Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents that contain intricate visual, textual, a…
ChuLo: Chunk-Level Key Information Representation for Long Document Understanding
Yan Li, Soyeon Caren Han, Yue Dai +1
Transformer-based models have achieved remarkable success in various Natural Language Processing (NLP) tasks, yet their ability to handle long documents is constrained by computati…
Graph-Based Multimodal Contrastive Learning for Chart Question Answering
Yue Dai, Soyeon Caren Han, Wei Liu
Chart question answering (ChartQA) is challenged by the heterogeneous composition of chart elements and the subtle data patterns they encode. This work introduces a novel joint mul…
Enhancing Document Key Information Localization Through Data Augmentation
Yue Dai
The Visually Rich Form Document Intelligence and Understanding (VRDIU) Track B focuses on the localization of key information in document images. The goal is to develop a method ca…