activity
20242026
collaborators

5 papers

cs.AI2026

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

Lewei Xu, Yihao Ding, Zihan Xu +5

Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…

cs.CV2026

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

Yihao Ding, Siwen Luo, Yue Dai +6

Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents that contain intricate visual, textual, a…

cs.AI2026

Docs2Synth: A Synthetic Data Trained Retriever Framework for Scanned Visually Rich Documents Understanding

Yihao Ding, Qiang Sun, Puzhen Wu +3

Document understanding (VRDU) in regulated domains is particularly challenging, since scanned documents often contain sensitive, evolving, and domain specific knowledge. This leads…

cs.CL2025

MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering

Shuo Yang, Siwen Luo, Soyeon Caren Han +1

Visual Question Answering (VQA) requires reasoning across visual and textual modalities, yet Large Vision-Language Models (LVLMs) often lack integrated commonsense knowledge, limit…

cs.CL2024

Multimodal Commonsense Knowledge Distillation for Visual Question Answering

Shuo Yang, Siwen Luo, Soyeon Caren Han

Existing Multimodal Large Language Models (MLLMs) and Visual Language Pretrained Models (VLPMs) have shown remarkable performances in the general Visual Question Answering (VQA). H…