collaborators

7 papers

cs.AI2026

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution

Lewei Xu, Yihao Ding, Zihan Xu +5

Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds a model's context window. Prior work has…

cs.LG2026

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth

Zihan Xu, Puzhen Wu, Lawrence Chun Man Lau +4

Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned ima…

cs.CV2026

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

Yihao Ding, Siwen Luo, Yue Dai +6

Visually Rich Document Understanding (VRDU) has become a pivotal area of research, driven by the need to automatically interpret documents that contain intricate visual, textual, a…

cs.AI2026

MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation

Yi Lin, Yihao Ding, Yonghui Wu +1

Automated 3D radiology report generation often suffers from clinical hallucinations and a lack of the iterative verification found in human practice. While recent Vision-Language M…

cs.CV2026

CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification

Hexin Dong, Yi Lin, Pengyu Zhou +25

Chest X-ray (CXR) interpretation is hindered by the long-tailed distribution of pathologies and the open-world nature of clinical environments. Existing benchmarks often rely on cl…

cs.CV2025

A Disease-Aware Dual-Stage Framework for Chest X-ray Report Generation

Puzhen Wu, Hexin Dong, Yi Lin +2

Radiology report generation from chest X-rays is an important task in artificial intelligence with the potential to greatly reduce radiologists' workload and shorten patient wait t…